jeffy-build
Retrain Jeffy classifier heads from Hugging Face datasets
TLDR
SYNOPSIS
jeffy-build [--datasets name ...] [--out dir] [--C value] [--max-train n] [--max-test n]
DESCRIPTION
jeffy-build is a console script from the jeffy-classify Python package. It downloads the source Hugging Face datasets for Jeffy's text heads, embeds them with BAAI/bge-large-en-v1.5, fits a scaler and logistic regression per task, writes artifacts with integrity hashes, and prints train/test accuracy.The built-in dataset list is banking77, clinc_oos, massive_intent, ag_news, dbpedia, sst2, emotion, imdb, sms_spam, snli, tweet_eval_sentiment, tweet_eval_emotion, and tweet_eval_offensive. doom_fire is not in this list (it is a numeric-feature head, not a Hugging Face text dataset).Each artifact lands under `--out`/taskid/ with a `manifest.json`. A pack-level `packmanifest.json` records encoder identity, classifier settings, and per-dataset scores. The README times a full rebuild at about 40 minutes and ~5 GB of dataset downloads.This command needs the optional build extra (`datasets`) on top of the core package: `pip install 'jeffy-classify[build]'` or `uv pip install -e ".[build]"` from a clone.
PARAMETERS
--datasets name ...
Dataset ids to train. Default is every key in the built-in table. Unknown names are printed and skipped.--out dir
Output directory. Default `data/model_pack`.--C value
Logistic-regression inverse regularization. Default `0.01`.--max-train n
Maximum training examples per dataset (random subsample, seed 42). Default `10000`.--max-test n
Maximum test examples per dataset (random subsample, seed 42). Default `2000`.
CAVEATS
A full run downloads several gigabytes from Hugging Face and loads the 1.2 GB encoder. Network, disk, and RAM requirements are those of sentence-transformers plus the datasets library.sst2 evaluates on the `validation` split (official test labels are not public). sms_spam uses a random 80/20 split. snli drops unlabeled (`-1`) rows and joins premise and hypothesis with ` [SEP] `. massive_intent is English only.Dataset licenses vary (CC BY, CC BY-SA, academic, Twitter TOS). Redistribution of retrained heads is not independently cleared for every source. See `ATTRIBUTION.md`.This rebuilds text heads only. It does not produce doom_fire.
HISTORY
jeffy-build shipped with Jeffy on GitHub in October 2026 (package jeffy-classify, MIT license, author Nico Brenner).
SEE ALSO
jeffy-serve(1), jeffy-evaluate(1), jeffy-train(1), python(1), uv(1), pip(1), hf(1)
