Do LLM benchmarks measure real knowledge or just instruction-following skills?
Models are frequently penalized on benchmarks for format errors, obscuring their true capabilities.
Read on to learn how we introduce soft-prompt tuning as a fairer way to evaluate LLMs. 🧵
We create specialized large language models for a sovereign Europe. Join us: jobs.ashbyhq.com/AlephAlpha #artificialintelligence, #writtenbyahuman
- The velocity of a team depends on tooling. Our model factory pioneers Model Training as Code, letting us scale our hill-climbing efforts while maintaining acceleration. Read about it below, and stay tuned for more: aleph-alpha.com/en/blog/model-…
- Today, we announce a landmark agreement with @cohere. By uniting our European research depth with global AI scale, we are building a transatlantic AI powerhouse to give enterprises control over their AI. Learn more about our shared vision here: businesswire.com/news/home/2026…
- We are releasing a detailed tech report for the tokenizer-free model family first introduced at ICLR 2025. These models achieve a significantly higher compression rate than standard LLMs, especially beyond English. 🧵
- Introducing Alpha-MoE: A fused megakernel for faster tensor parallel inference. With up to 200% speedups for MoE models in TP deployments. Optimized for Hopper. Built for sovereignty and scale. aleph-alpha.com/alpha-moe-a-me…

