Kev 1.0: Open-Weight Family of Decision Models (Kev-27B, Kev-9B, Kev-4B, Kev-0.8B)
- Type
- model
- Venue
- Jared Palmer blog ('Introducing Kev'), 1 Oct 2026; open-weight release
- Year
- 2026
- Source
- blog
- Access
- open
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Kev 1.0 is a family of four open-weight (Apache-2.0) 'decision models' — Kev-27B (new), Kev-9B (updated), Kev-4B, and Kev-0.8B — built on Qwen3.5/3.8 backbones. Each model adds a pointer head that scores user-supplied answer options (yes/no, multiple choice, ordered scales) and returns calibrated probabilities per answer, instead of generating text: the document is processed once and cached, and every question branches off that single read via <decide> tokens. The interface matches TypeSafe's Jev API, so TypeSafe SDK apps can switch by changing endpoint and model name. Kev-27B is a full fine-tune of the Qwen3.8-27B post-trained text backbone on ~146k records / 337k questions (documents up to 32k tokens; 64k-token document window), blending 85% of the new fine-tune with 15% of the previous prerelease checkpoint; the smaller models use Qwen3.5 backbones with LoRA adapters. Weights are on Hugging Face, code and a GPU/Apple-Silicon (MLX) server are on GitHub, and agent skills (kev-finetune, kev-deploy) package fine-tuning on your own data and deployment to Modal.
Keywords
Kev · decision models · Jared Palmer · TypeSafe · Jev · Qwen3.8 · Qwen3.5 · pointer head · text classification · calibrated probabilities · LoRA · MLX · kev-finetune · kev-deploy
Topics
decision models, text classification, open weights, Qwen fine-tuning, calibrated probabilities, pointer head, fine-tuning tooling
Research notes
- Discovery: @jaredpalmer X thread (5 parts) 2026-10-01 (https://x.com/jaredpalmer/status/2105803379587068011)
- Weights (Apache-2.0): https://huggingface.co/collections/jaredpalmer/kev (jaredpalmer/kev-27b, kev-9b, kev-4b, kev-0.8b; also a kev-0.5b v0.1 prototype on Qwen2.5-0.5B)
- Code: https://github.com/jaredpalmer/kev (server runs on GPUs and Apple Silicon via MLX; Kev-4B tryable in browser via HF ZeroGPU Space)
- How it works: pointer head scores supplied options without generating answer tokens; document prefix computed once and cached; each <decide> token scores options, softmax -> probabilities, calling code sets the threshold; Kev only scores options you supply — no explanations, no retrieval of missing facts; answer order can shift probabilities slightly
- Training Kev-27B: full fine-tune of Qwen3.8-27B (post-trained) text backbone on ~146k examples / 337k questions, docs up to 32k tokens; corpus = existing Kev data + licensed public tasks + synthetic decisions written by open-weight models + code-generated documents/agent logs with exact answers; added tone, grounding, prompt-injection and personal-data-detection examples; records screened against frozen eval sets for exact/near matches; no Jev outputs used; sharing document computation across a document's questions made training 2-3x faster; one full run ~16h on 8x H200 (~$650 GPU time); released backbone = 85% new fine-tune + 15% previous Kev-27B prerelease checkpoint; training corpus not public
- Evaluation (author's own harness, not an independent leaderboard; selection not fully blind; Qwen base exposure unknown): Generalization (14 public datasets, 3,089 test questions excluded from fine-tuning): Kev-27B 75.7%, Kev-9B 69.8%, Kev-4B 69.0%, Kev-0.8B 58.6%; chance-corrected generalization index 52.3 / 41.0 / 38.0 / 23.3
- Short-text decisions (656 dev): 85.1 / 82.0 / 81.7 / 64.8
- Knowledge and evidence (1,046 dev): 82.0 / 78.0 / 77.4 / 58.6
- Policy and rule reasoning (held-out, 1,088 test): 91.8 / 83.4 / 80.3 / 66.5
- Sorting consumer complaints (936): 90.8 / 90.0 / 90.3 / 85.1
- Developer-tool decisions (1,071): 79.0 / 79.1 / 75.6 / 63.7 (Kev-9B edges out 27B)
- Calibration: pooled calibration error 0.019 (27B), 0.034 (9B), 0.029 (4B), 0.042 (0.8B); threshold targeting 5% dev error -> 27B accepts 54% of questions at 5.5% observed error; temperature fitted on held-out data ships with checkpoint
- Serving (Modal, Oct 2026 prices): Kev-27B on H200 29 req/s, ~$44.09/1M requests, 67ms median; Kev-9B on H100 80 req/s, $13.80, 24ms; Kev-4B on L40S or 32GB Mac 51 req/s, $10.54, 42ms; Kev-0.8B on L4 $3.54, 23ms; Kev-27B reads a new 64k-token document in ~9.4s model time on H200, subsequent questions ~0.7s cached; locally, 0.8B/4B run on Apple Silicon via MLX (32GB M5: Kev-4B 5 questions on short text in 721ms, 136ms cached)
- Kev-9B update: one extra fine-tuning pass (~3h, one GPU) — policy/rule 58%->83%, dev-tool 64%->79%, complaints 83%->90%; previous version at jaredpalmer/kev-9b@v1
- Fine-tuning: kev-finetune skill (define questions, prepare/generate labeled data, fine-tune, fit temperature, compare vs starting checkpoint) and kev-deploy skill (Modal behind authenticated HTTPS, scales to zero, ~35s cold start); both on skills.sh, work with Claude Code/Codex/Devin; small Kev-4B fine-tune ~$1 GPU; example: one pass over 5,219 labeled complaints improved held-out accuracy 80%->90%
- Limitation noted by author: new Kev-27B is less accurate and more overconfident on long legal contracts than the previous prerelease; for that workload compare jaredpalmer/kev-27b@v1-lora
- License: Apache-2.0