jeffy-evaluate
Evaluate frozen Jeffy classifier heads on held-out test sets
TLDR
SYNOPSIS
jeffy-evaluate [--tasks name ...] [--pack-dir dir] [--out dir] [--max-test n] [--baselines] [--latency] [--device cpu|mps|cuda]
DESCRIPTION
jeffy-evaluate is a console script from the jeffy-classify Python package. It loads frozen artifacts from a model pack (no retraining), scores them on the same Hugging Face splits jeffy-build used, writes per-example predictions, and prints a summary table.Default `--pack-dir` is `data/modelpack` and default `--out` is `data/evalresults`. Output includes one `{task}predictions.jsonl` per task plus `benchmark.json` (scores, optional baselines and latency, hardware info). Accuracy is reported with a bootstrap 95% confidence interval. For **clincoos** the table also splits in-scope versus out-of-scope accuracy.`--baselines` fits a majority-class dummy and a 3-fold CV-tuned TF-IDF plus logistic regression on the training split. `--latency` times single-example inference after a short warmup. Thread counts are pinned (`torch` threads and `OMPNUMTHREADS`/`MKLNUMTHREADS`) so latency numbers are comparable.Tasks without a dataset config are skipped. That includes doom_fire. Like jeffy-build, this command needs the build extra (`datasets`) to download evaluation splits: `pip install 'jeffy-classify[build]'`.
PARAMETERS
--tasks name ...
Task ids to score. Default is every capability loaded from the pack that also has a dataset config.--pack-dir dir
Model pack to load. Default `data/model_pack`.--out dir
Directory for `benchmark.json` and per-task prediction JSONL. Default `data/eval_results`.--max-test n
Maximum test examples per task (seed 42). Default `2000`.--baselines
Run majority-class and tuned TF-IDF+LR baselines on the training split.--latency
Measure per-example total, embedding, and classifier latency.--device cpu|mps|cuda
Encoder device. Default `cpu`.
CAVEATS
This does not retrain heads. Scores should match jeffy-build test accuracy within about `0.001`; larger gaps are printed as a discrepancy.`--baselines` is slow: the TF-IDF grid covers analyzers, n-grams, feature caps, and several `C` values with 3-fold CV.sst2 uses the `validation` split. sms_spam uses a random split. snli filters unlabeled rows. Hugging Face downloads and the 1.2 GB encoder apply here the same as for jeffy-build.
HISTORY
jeffy-evaluate shipped with Jeffy on GitHub in October 2026 (package jeffy-classify, MIT license, author Nico Brenner).
SEE ALSO
jeffy-build(1), jeffy-serve(1), jeffy-train(1), python(1), uv(1), pip(1), hf(1)
