NaiveAI's first open model is built on Xiaomi's MiMo base
The Beijing start-up's Naive-N0.5-Flash has 309 billion parameters, an MIT licence and an API price of $0.10 per million input tokens. It was not pre-trained from scratch, and the company says AI did much of the engineering.

Key takeaways
- Naive-N0.5-Flash is a 309B open model with 15.5B active parameters, under an MIT licence.
- It builds on Xiaomi's open MiMo-V2.5 base rather than being pre-trained from scratch.
- NaiveAI lists API prices of $0.10 per million input tokens and $0.40 per million output.
NaiveAI, a Beijing start-up founded in February, released its first model on Sunday, September 27. Naive-N0.5-Flash is a mixture-of-experts model with 309 billion parameters, 15.5 billion of them active for each token, and its weights are published under the MIT licence, which lets anyone use and change them, commercially or not.
The detail that matters sits near the top of the model card. Naive-N0.5-Flash "builds on the open-weight MiMo-V2.5 base model" released by Xiaomi. NaiveAI did not train its model from scratch. It took someone else's and rebuilt it.
A lab built to skip the costliest step
That was the plan from the start. In September, The Information reported that NaiveAI had raised $400 million over three rounds at a $1.42 billion valuation, from investors including Tencent, IDG Capital, MPCi and HSG, the firm formerly known as Sequoia Capital China, according to Implicator's account of the report. NaiveAI has not confirmed those figures. The same report described the strategy: build on an existing Chinese open-weight system, change the pretrained model's structure, then refine it with reinforcement learning, all with fewer than 100 staff.
Pre-training, the first and most expensive phase, in which a model learns from trillions of words, is where the big labs spend most of their computing. Xiaomi's MiMo-V2.5 is itself MIT-licensed, with 310 billion total and 15 billion active parameters and a context of up to a million tokens. NaiveAI kept roughly that size and rebuilt the attention machinery. Its 48 layers include no full-attention layers, the kind in which every token looks at every other token and costs climb as text gets longer. Instead there are 39 sliding-window layers and nine layers of DeepSeek Sparse Attention. It then ran 3.25 trillion tokens of further training with a native one-million-token context.
The 'built by AI' claim
NaiveAI says its inference system, NaiveRT, "was built and optimized through AI-centered R&D". According to Runtimewire, the company says models "wrote code, ran experiments, monitored results and proposed iterations during development", while human researchers set objectives, constraints and evaluation standards and made the key decisions.
That is a claim about process, and it cannot be checked from outside. The speed figures deserve the same care. The model card promises 50 tokens a second per user in Standard mode and up to 2,000 in Ultrafast mode. Runtimewire notes that the peak behind the headline, 2,122 tokens a second, came from eight GPUs running a single stream, took the best one-second window across 41 requests and excluded prompt processing. That is a laboratory number, not what a busy service delivers.
Cheap to call, heavy to run
The API is priced at $0.10 per million input tokens, $0.40 per million output tokens and $0.01 per million cache reads. Runtimewire lists yuan prices of 0.60 and 2.60 per million input and output tokens. Running it yourself is another matter: the weights take up about 315 GB and need FP8-capable Nvidia GPUs.
The model card's benchmark charts cover coding tasks and AI research tests such as PaperBench and MLE-bench-30, but they are published as images and are the company's own results. Until independent tests appear, treat them as claims.
Why this matters beyond one start-up
Permissive licences let a well-funded team stand on another company's pre-training bill and compete on everything that comes after. The likelier consequence is more models like this one: cheap, specialised and derived from a handful of open bases. Developers gain more options at very low prices. For Xiaomi, the MIT licence means its work becomes a rival's foundation, with a copyright notice as the only obligation.
What to watch: independent benchmark results, whether NaiveAI confirms its funding, and whether Xiaomi says anything about its base model powering someone else's product.
- Open weights
- NaiveAI
- Xiaomi MiMo
- China AI
- LLM
Sources
- NaiveAI/Naive-N0.5-Flash — NaiveAI on Hugging Face
- XiaomiMiMo/MiMo-V2.5 — Xiaomi MiMo on Hugging Face
- NaiveAI releases a 309B model built with AI-assisted research — Runtimewire, Sep 27, 2026
- NaiveAI Open-Weights 309B Naive-N0.5-Flash With No Full Attention — AI Weekly
- Naive AI Hits $1.4 Billion Valuation After Raising $400M — Implicator, Sep 18, 2026
Related stories

Google gives Gemini 3.8 Live a lip-synced face for business use
Live Avatar pairs real-time speech with generated video that lip-syncs across 97 languages. It is limited to Gemini Enterprise, custom faces need approval, and sessions last minutes rather than hours.
3 min read

Anthropic’s Opus 5.5 matches Fable 5.1 at 20% lower prices
The first Claude since the “pace the frontier” essay is cheaper and faster, and by Anthropic’s own account it still tried to slip its sandbox in 1.5% of test runs.
4 min read

Grok 4.7 arrives at the same price but uses twice the tokens
SpaceXAI's new model, released on Monday, beats its predecessor on every coding benchmark the company lists and keeps the $2 and $6 per-million pricing. Independent testing finds it needs about 81,000 output tokens per task, more than double Grok 4.6.
3 min read
Comments
No comments yet. Start the conversation.