LinuxCommandLibrary
GitHubF-DroidGoogle Play Store

jeffy-train

Train a custom Jeffy text classifier from CSV or JSONL

TLDR

Train the bundled reviews example and save the head
$ jeffy-train --example --save-dir [my_models]
copy
Train from a CSV (text and label columns)
$ jeffy-train --input [data.csv] --text-col [text] --label-col [label] --task-id [reviews] --save-dir [my_models]
copy
Train from JSONL
$ jeffy-train --input [data.jsonl] --text-col [text] --label-col [label] --task-id [reviews] --save-dir [my_models]
copy
Train from a TSV
$ jeffy-train --input [data.tsv] --text-col [text] --label-col [label] --task-id [reviews] --save-dir [my_models]
copy
Change regularization (smaller is more conservative)
$ jeffy-train --example --C [0.001] --save-dir [my_models]
copy
Serve the saved head
$ JEFFY_PACK_DIR=[my_models] jeffy-serve
copy

SYNOPSIS

jeffy-train (--input file | --example) [--text-col name] [--label-col name] [--task-id id] [--C value] [--save-dir dir] [--test-size frac]

DESCRIPTION

jeffy-train is a console script from the jeffy-classify Python package. It embeds labeled texts with BAAI/bge-large-en-v1.5, fits a StandardScaler plus LogisticRegression, prints train/test scores, and optionally writes a model-pack artifact that jeffy-serve can load.Input is CSV, TSV, or JSONL. Column names default to `text` and `label`; `--text-col` and `--label-col` also accept the aliases `--text-key` and `--label-key`. `--example` trains on the 24-row product-review CSV bundled in the package and sets `--task-id` to `reviews` when you leave the default `custom`.After training, the process always enters an interactive prompt (`>`) that classifies typed lines until Ctrl-C or EOF. Use `--save-dir` so the head is written before that loop. Serving a saved pack is `JEFFYPACKDIR=dir jeffy-serve`.The same library exposes `jeffy.train.trainclassifier` and `trainfeature_classifier` for Python callers. Numeric feature training is not wired to this CLI; the command only reads text columns.

PARAMETERS

--input file

Path to `.csv`, `.tsv`, or `.jsonl`. Required unless --example is set.
--example
Train from the bundled `reviews.csv` (24 rows, two classes). Sets --task-id to `reviews` when it is still `custom`.
--text-col, --text-key name
Text column or JSON key. Default `text`.
--label-col, --label-key name
Label column or JSON key. Default `label`.
--task-id id
Name written into the artifact. Default `custom`.
--C value
Logistic-regression inverse regularization. Default `0.01`. Lower values keep predictions closer to uniform; higher values fit individual examples more tightly.
--save-dir dir
Directory to write `dir/task-id/` (scaler, classifier, `manifest.json`). Omit to keep the model only in memory for the interactive prompt.
--test-size frac
Hold-out fraction. Default `0.2`. Applied only when there are at least 10 examples. The Python API also supports `cv_folds`; this CLI does not expose that flag (the trainer still runs 3-fold CV when the train split is large enough).

CAVEATS

Needs at least two examples. The first run downloads the encoder (about 1.2 GB). Encoding runs on CPU.The interactive test loop always starts after a successful train, including when `--save-dir` was used. Redirecting stdin or piping a file into the command will feed that loop.Custom artifacts may include a joblib pickle. Load them only from trusted sources. The CLI does not train feature-based heads such as doom_fire.

HISTORY

jeffy-train shipped with Jeffy on GitHub in October 2026 (package jeffy-classify, MIT license, author Nico Brenner).

SEE ALSO

jeffy-serve(1), jeffy-build(1), jeffy-evaluate(1), python(1), uv(1), uvx(1), pip(1)

RESOURCES

Copied to clipboard
Kai