Skip to content

Model Card: AliceAI-Foundation-80B-A3B-Base

Sep 28, 20261 min
TL;DRAliceAI-Foundation-80B-A3B-Base: released by Yandex on 2026-09-21, trained fully from scratch (not a Qwen/Llama fine-tune); 80B total / 3B active MoE parameters, 262,144-token context window; Apache-2.0, no official hosted API pricing yet; scores 91.1 on MATH-500, beating similarly-sized Qwen3.5-35B-A3B-Base and the much larger DeepSeek-V4-Flash-Base (284B-A13B); leads every comparison on the Russian-language WikiWebFacts benchmark at 86.5; it's a raw base model with no SFT/RL alignment, so it can't be used as an agent out of the box

🌏 中文版

Model Info

ItemValue
Model IDyandex/AliceAI-Foundation-80B-A3B-Base
VendorYandex
Parameters80B total / 3B active (MoE, 512 routed experts + 1 shared expert, top-10 routing)
Context Window262,144 tokens (~256K)
Input Pricing (USD/1M tokens)Not available (fully open-source base model; no official hosted API, no third-party inference provider yet)
Output Pricing (USD/1M tokens)Not available (same as above)
Open SourceYes (Apache-2.0)
Release Date2026-09-21
Official AnnouncementYandex Investor Relations
HuggingFaceyandex/AliceAI-Foundation-80B-A3B-Base
FamilyAlice AI (Yandex's assistant model family; this release is an experimental step toward a future unified reasoning model)

Highlights

  • Trained entirely from scratch — not a fine-tune of Qwen or Llama. The 48-layer network repeats a 4-layer pattern (three KDA linear-attention blocks followed by one Gated Attention block), with a 512-expert MoE attached to every layer; only 3.75% of parameters activate per token, sharply cutting inference compute
  • Strong math and competition reasoning: MATH-500 hits 91.1, ahead of the similarly-sized Qwen3.5-35B-A3B-Base (81.9) and the 3.5x-larger DeepSeek-V4-Flash-Base (284B-A13B, 80.7); AIME 2026 pass@32 ties Qwen3.5-35B-A3B-Base for first place at 96.7
  • Dominant on Russian factual knowledge: scores 86.5 on Yandex's own WikiWebFacts benchmark, 24.1pp ahead of Qwen3.5-35B-A3B-Base's 62.4, and matches Yandex's own previous closed model Alice AI LLM despite that model having 3x the parameters and 7x the active parameters
  • On the engineering side, Yandex rebuilt its training optimizer so GPU-to-GPU data transfer overlaps with computation, roughly doubling training step speed; a cascade of classifiers now pre-filters documents before the expensive scoring model runs, cutting what would have been 200,000+ GPU-hours of compute by more than 10x while keeping about 95% of useful documents

Benchmark Results

BenchmarkAliceAI-Foundation-80B-A3B-BasePrevious GenBest Competitor
MATH-500 (5-shot)91.1N/A (Yandex's first from-scratch open base model)Qwen3.5-35B-A3B-Base 81.9
AIME 2026 pass@3296.7N/AQwen3.5-35B-A3B-Base 96.7 (tied)
WikiWebFacts (Russian factual knowledge, 5-shot)86.5Alice AI LLM (Yandex's previous closed model, comparable score despite 3x the parameters and 7x the active parameters)DeepSeek-V4-Flash-Base 83.2
LiveCodeBench v5-6 CoT 1-shot pass@150.5N/AQwen3.5-35B-A3B-Base 50.4
CodeForces C++ pass@868.9N/AQwen3.5-35B-A3B-Base 73.7 (highest overall)

⚠️ All figures above are Yandex's own measurements on an internal evaluation setup (vLLM, t=0 or t=1) and await independent reproduction. WikiWebFacts, HardMultiQA, CultCat, and the EduBench series are Yandex's own Russian-language benchmarks, released publicly with full evaluation protocols so others can reproduce them.

Comparison with Previous Gen and Competitors

This is Yandex's first fully from-scratch, publicly released open base model — not a fine-tune built on top of Qwen or Llama. The training corpus, architecture, and hyperparameters were all designed from the ground up. Yandex positions it as an experimental step toward a future unified reasoning model that will power Alice AI's agentic capabilities, letting users delegate specific actions to the assistant.

Against Qwen3.5-35B-A3B-Base, which has a comparable number of active parameters, the two trade wins: AliceAI leads MATH-500 by 9.2pp and ties on AIME 2026, but trails by 4.8pp on CodeForces C++ competitive programming. The real differentiator is Russian-language and domain-expert knowledge — WikiWebFacts, EduBench History/Literature, and the law and medicine tracks of ExpertFactsQA all go to AliceAI, often by more than 20pp, beating even DeepSeek-V4-Flash-Base despite that model having several times the parameters and active parameters.

There's no real basis for a pricing comparison: AliceAI ships fully open under Apache-2.0 with no official hosted API, so running it means self-hosting via vLLM or Transformers (the full BF16 checkpoint is about 163GB). That's a different approach from most closed-source flagships, which typically pair a release with API pricing on day one — a sign that Yandex's current goal is validating the architecture and training methodology rather than chasing market share.

What This Means for Agent Development

A MoE architecture that mixes linear attention (KDA) with standard attention while keeping active parameters down to 3.75% is a useful reference point for the cost/context tradeoffs in agent architecture design.

  • If you're experimenting with agent inference architectures: this is one of the few publicly released large models that interleaves linear attention (KDA) with standard Gated Attention layers on top of an extremely fine-grained 512-expert MoE — its cost structure, pairing a 256K context window with only 3B active parameters, is worth studying
  • If your product targets Russian-speaking or multilingual markets: this is currently one of the strongest open models for Russian factual and domain-expert knowledge (law, medicine, education), and Yandex has released two Russian-language benchmarks, WikiWebFacts and HardMultiQA, so you can verify that yourself
  • Not a fit if: this is a raw pretrained base model with no instruction tuning or RL alignment, so it can't be used directly as a chatbot or for agent action decisions — you'll need to do your own SFT/RL. There's also no official hosted API yet, meaning you'll need to self-host a vLLM deployment (a 163GB BF16 checkpoint), which is not a low deployment bar

Today's Takeaway

It's easy to assume that any open base model competitive with much larger models is a fine-tune built on Qwen or Llama. AliceAI-Foundation is a counterexample: a search-engine company designed a hybrid attention architecture from scratch, without an obviously overwhelming data advantage, and still beat derivative models several times its size on math reasoning and domain-specific knowledge. That suggests the "base model moat" comes more from training methodology and data-curation precision than from raw parameter count alone.

References