BTL-4
- Type
- model
- Venue
- Bad Theory Labs
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:50:00Z
- Verified
- 2026-08-14T19:50:00Z
Summary
Bad Theory Labs BTL-4: 35B Qwen3.5 MoE (~2.1B active) finetuned from Ornith-1.0-35B on trajectories whose code actually ran and passed tests. Vendor: SWE-bench Verified 78.4%; BFCL v4 AST 73.5% (+4.3 over base, paired, 1240 cases); LiveCodeBench v6 66.1% (easy 99.1 / medium 86.7 / hard 60.5; 442 problems; 16K→32K output budget +5.2). Native 262K context. Compact is one 9.96 GB GGUF at 2.30 bpw for a 16 GB card; tweet claims KV ~20 KB/token so 262K still fits. Per-group clip search (12 s, no calib data) moved 2-bit behavioural retention 77.1%→95.8%; Compact reproduced 111/118 full-precision gate behaviours (94.1%). Stock llama.cpp/Ollama/LM Studio. Apache-2.0. Companion Macaw 2.7B is a separate Mac agent, not this row.
Keywords
btl-4 · bad-theory-labs · moe · quantization · swe-bench · llama-cpp · x
Topics
open-weight MoE, agentic coding, extreme quantization
Research notes
- Primary: HF model card (Apache-2.0). Discord/X https://x.com/Badtheorylabs/status/2085359932900082039 via fxtwitter (note tweet quoting the 2026-08-05 BTL-4/Macaw launch). Compact https://huggingface.co/badtheorylabs/BTL-4-Compact. Macaw https://github.com/Badtheorylabs/Macaw. Vendor benches. Open weights, not a new hosted corpus, so no datasets_local row.