← Back to explorer

Apriel-1.5-15B-Thinker: Mid-training is all you need

Type
other
Venue
arXiv / ServiceNow

Summary

Starts from Pixtral-12B, depth-upscales the decoder 40→48 layers, then two-stage multimodal CPT (foundational text/vision then synthetic visual reasoning) and high-signal text SFT with reasoning traces. No RL or preference optimization. Artificial Analysis Intelligence Index 52 (matches DeepSeek-R1-0528); AIME 2025 87.5, IFBench 61.7, τ²-Bench Telecom 68.4. Ten image benchmarks average within ~5 points of Gemini-2.5-Flash / Claude Sonnet 3.7. Weights MIT at ServiceNow-AI/Apriel-1.5-15b-Thinker.

Keywords

apriel · mid-training · pixtral · multimodal · servicenow · compact-llm · sft · no-rl

Topics

multimodal reasoning, mid-training, compact LLMs

Research notes

  • Primary: arxiv abs (cs.AI). Paper states model/recipes/evals released under MIT; HF cardData.license mit. ServiceNow SLAM lab. Weights https://huggingface.co/ServiceNow-AI/Apriel-1.5-15b-Thinker (473 likes / 222 downloads at check; usedStorage ~29.7 GB). HF paper page 125 upvotes; linked models not copied into hf_* fields. Discord posted PDF. Training mix is not a standalone public corpus on abs, so no datasets_local row. License field left blank per catalog convention.