LinuxCommandLibrary
GitHubF-DroidGoogle Play Store

jeffy-evaluate

Evaluate frozen Jeffy classifier heads on held-out test sets

TLDR

Evaluate every text head in a model pack
$ jeffy-evaluate --pack-dir [data/model_pack] --out [data/eval_results]
copy
Evaluate named tasks only
$ jeffy-evaluate --tasks [banking77] [ag_news] --pack-dir [data/model_pack]
copy
Also run majority-class and TF-IDF baselines
$ jeffy-evaluate --baselines --pack-dir [data/model_pack]
copy
Also measure single-example latency
$ jeffy-evaluate --latency --device [cpu] --pack-dir [data/model_pack]
copy
Cap test-set size (default 2000)
$ jeffy-evaluate --max-test [2000] --pack-dir [data/model_pack]
copy

SYNOPSIS

jeffy-evaluate [--tasks name ...] [--pack-dir dir] [--out dir] [--max-test n] [--baselines] [--latency] [--device cpu|mps|cuda]

DESCRIPTION

jeffy-evaluate is a console script from the jeffy-classify Python package. It loads frozen artifacts from a model pack (no retraining), scores them on the same Hugging Face splits jeffy-build used, writes per-example predictions, and prints a summary table.Default `--pack-dir` is `data/modelpack` and default `--out` is `data/evalresults`. Output includes one `{task}predictions.jsonl` per task plus `benchmark.json` (scores, optional baselines and latency, hardware info). Accuracy is reported with a bootstrap 95% confidence interval. For **clincoos** the table also splits in-scope versus out-of-scope accuracy.`--baselines` fits a majority-class dummy and a 3-fold CV-tuned TF-IDF plus logistic regression on the training split. `--latency` times single-example inference after a short warmup. Thread counts are pinned (`torch` threads and `OMPNUMTHREADS`/`MKLNUMTHREADS`) so latency numbers are comparable.Tasks without a dataset config are skipped. That includes doom_fire. Like jeffy-build, this command needs the build extra (`datasets`) to download evaluation splits: `pip install 'jeffy-classify[build]'`.

PARAMETERS

--tasks name ...

Task ids to score. Default is every capability loaded from the pack that also has a dataset config.
--pack-dir dir
Model pack to load. Default `data/model_pack`.
--out dir
Directory for `benchmark.json` and per-task prediction JSONL. Default `data/eval_results`.
--max-test n
Maximum test examples per task (seed 42). Default `2000`.
--baselines
Run majority-class and tuned TF-IDF+LR baselines on the training split.
--latency
Measure per-example total, embedding, and classifier latency.
--device cpu|mps|cuda
Encoder device. Default `cpu`.

CAVEATS

This does not retrain heads. Scores should match jeffy-build test accuracy within about `0.001`; larger gaps are printed as a discrepancy.`--baselines` is slow: the TF-IDF grid covers analyzers, n-grams, feature caps, and several `C` values with 3-fold CV.sst2 uses the `validation` split. sms_spam uses a random split. snli filters unlabeled rows. Hugging Face downloads and the 1.2 GB encoder apply here the same as for jeffy-build.

HISTORY

jeffy-evaluate shipped with Jeffy on GitHub in October 2026 (package jeffy-classify, MIT license, author Nico Brenner).

SEE ALSO

jeffy-build(1), jeffy-serve(1), jeffy-train(1), python(1), uv(1), pip(1), hf(1)

RESOURCES

Copied to clipboard
Kai