Sber Releases GigaChat 3.5 Reasoning Weights and Inference Code
Sber released GigaChat 3.5 Reasoning on September 10, presenting it as the first GigaChat model with full reasoning capabilities. In a company-authored technical post, Sber said it had published the model’s weights and the code needed to run it on Hugging Face. The model is also available to try at giga.chat. Developers can choose FP8 weights for inference or BF16 weights for fine-tuning and custom quantization.
The company reports substantial gains over GigaChat 3.5 Instant in three evaluations. GPQA-Diamond increased from 61.11 to 82.32, AIME-2026 mean@32 from 67 to 92, and IFBench Loose Prompt from 43.66 to 77.00. These figures come from Sber’s own evaluation configurations, and the source provides no independent reproduction.
Rather than applying reinforcement learning to a single model across every domain, Sber started with a supervised fine-tuning checkpoint and developed six separate domain models. Each was trained independently with online reinforcement learning and a domain-specific reward system. The six specialists were subsequently consolidated into one model through on-policy distillation.
The specialists addressed STEM problems, conventional coding, repository-level coding, general agent tasks, user dialogue, and instruction following. Verification depended on the task: final answers could be checked against references, code could be executed, repository patches could undergo tests, and agent runs could be judged by the environment’s final state. Sber says this structure avoided competition between domains for the same weights and enabled more precise reward design.
For reinforcement learning, Sber used CISPO, an algorithm related to GRPO. The company explains that CISPO restricts the contribution of tokens whose probabilities have moved too far instead of removing their training signal. This is intended to retain signals around uncommon reasoning turns, including detecting an error or revisiting an earlier step. Sber says CISPO converged faster than GRPO when training its specialists.
Sber also adjusted the training set as checkpoints improved. Before each stage, the current checkpoint was evaluated across the task pool, and tasks solved in more than 75% of attempts were excluded. According to the company, eliminating rollouts for examples the model had already mastered cut inference compute during training by 50%. An adaptive length penalty discouraged unnecessary reasoning more strongly on easier tasks than on difficult ones.
Practical context: The two weight repositories support either direct inference or further adaptation, but performance in other settings still requires separate verification. For this release, the repository specialist was trained only through mini-SWE-agent, while most conventional programming tasks were in Python. The reported results therefore do not establish performance with other agent harnesses or programming languages.
| Specialist | Main tasks | Result verification |
|---|---|---|
| STEM | Mathematics, olympiad problems, natural sciences | Compare the final answer with a reference |
| Code | Algorithms, code edits, test generation | Execute the code |
| Code Agent | Tasks involving a real repository | Run tests after applying the patch |
| General Agent | Function calling, user dialogue, memory, search | Check the environment’s final state |
| Social Interaction | User dialogue | Pairwise assessment by an LLM judge |
| Instruction Following | Instruction following, formats, long context, structured output | Compare the final answer with a reference |
| Repository | Purpose |
|---|---|
| ai-sage/GigaChat3.5-432B-A28B-Reasoning | FP8 inference |
| ai-sage/GigaChat3.5-432B-A28B-Reasoning-bf16 | Fine-tuning and custom quantization |
Sources
Event date: 2026-09-10. Primary source date: 2026-09-10.