The inference model “Darwin-180B-RSI,” developed by AI deep-tech company Bidraft (CEO Kim Min-sik), has reached the top position across all five certified benchmark leaderboards on global AI platform Hugging Face. Bidraft announced on the 28th that it had achieved these results and stated that it has released the entire model as open source so that anyone can verify it.
This evaluation is significant in that it reflects performance in key categories officially reviewed by Hugging Face, rather than on metrics arbitrarily proposed by the company. In particular, the model achieved a perfect score of 100% on “AIME 2026” and “HMMT 2026,” both gateways to high-difficulty mathematics Olympiads. In these categories, where roughly 40 global models competed, the previous best scores were 97.1% and 92.7%, respectively. This is the first time a perfect score has been recorded on Hugging Face’s official mathematics leaderboard.
In the “GPQA Diamond” category, which measures scientific reasoning ability at the level of PhD experts, the model also recorded a high accuracy rate of 94.44%, surpassing the previous leader “Kimi-K3 (93.5%).” Bidraft’s earlier model, “Darwin-397B-ZTC,” also ranked third on the same metric, allowing the company to simultaneously occupy the top tier. In addition, it ranked first in “MMLU-Pro (88.12%),” which evaluates a wide range of specialized knowledge, and in the multimodal exam “MMMU-Pro (79.48%).”
The five categories dominated this time cover verification stages regarded as among the most difficult in the AI field: PhD-level scientific topics, multiple-choice questions across 14 specialized domains, university-level evaluations requiring chart and video interpretation, and Olympiad-representative mathematics problems.
The combination of three proprietary core technologies led to the performance improvement. Using “Darwin” merge technology, which diagnoses the strengths and weaknesses of the expert structure inside the model and transplants only optimal components, the company significantly reduced computation costs while scaling the model up to 397 billion parameters. It also introduced a “Recursive Self-Improvement (RSI)” technique, in which the model verifies its own answers and repeatedly trains, thereby shortening inference length by 11% while maintaining inference accuracy. In addition, it applied “Zero-Token Confidence (ZTC)” technology, which measures the internal state prior to answer generation to block hallucinations.
Bidraft, which aims to provide customized AI production (AI foundry), is expanding the ecosystem, with cumulative downloads exceeding 2.5 million, buoyed by the popularity of its lightweight model “POCKET-35B.” The company is currently participating in a Naver Cloud consortium and carrying out a national project to build a security-specialized AI foundation model.
CEO Kim Min-sik of Bidraft said the company had proven its world-class capabilities with AI technology that learns autonomously and measures its own reliability, and added that it will continue to develop AI that evolves on its own without additional human intervention.
ⓒ dongA.com. All rights reserved. Reproduction, redistribution, or use for AI training prohibited.
Popular News