Jiyuan Ludong has delivered its first large language model achievement since founding. According to QbitAI, the company founded by Wang Yunhe, former director of Huawei Noah's Ark Lab and head of Pangu LLM, has released NeoHorse, its first Agent-Native model, with compute support and infrastructure optimization from Infinigence AI, plus algorithm and training-method research participation from teams at Tsinghua University and Peking University. NeoHorse-1 includes 4B and 9B versions, focused on a set of capabilities required by Agent workflows: calling tools, reading environmental feedback, detecting errors, adjusting paths, and completing tasks.
文章图片 2
The company has long emphasized multi-model collaboration. Why start training its own models? NeoHorse's answer suggests the direction has not changed—this model is more like writing accumulated multi-model execution experience into model parameters for the first time. A core judgment from the Jiyuan Ludong team is that "models will not consolidate into one" will remain a long-term structural reality of the AI industry: the more models there are and the finer the division of labor, the more necessary it becomes to have a system that answers "which model should be used for this step, when should it be upgraded, and who takes over after failure." They call this layer Routing Harness. Its open-source project OpenSquilla has already connected multiple models, using a unified interface for fine-grained routing, model switching, and multi-model collaboration.
文章图片 4
NeoHorse's data source is worth noting. Its core corpus combines Agent execution signals generated by Routing Harness with public data to build an Agent-oriented post-training data system. When an Agent completes a task in the Harness, it leaves a complete execution record: what capabilities the task requires, what execution decisions the system made, and what feedback the environment ultimately gave. For example, the router may initially judge that a task only needs an ordinary model, then after consecutive failures during execution upgrade to a stronger model to complete it. This trajectory contains far more information than a single failure. The team observed that failure and repair trajectories may provide even more information—they can tell the model where errors are likely and demonstrate how to adjust after errors occur.
文章图片 6
Feeding all logs directly into training will not naturally produce a stronger Agent. Every trajectory must pass structural checks and be evaluated for execution quality across six dimensions: whether it fulfills the user's goal, follows instructions, uses tools reasonably, supports conclusions with evidence, recovers after errors, and terminates at the appropriate time. The team also deliberately distinguishes "task ended" from "goal satisfied"—the model outputting "completed" does not mean the deliverable meets requirements. Beyond filtering, routing signals also play a role in training: the router estimates the capabilities a task requires and arranges sample order accordingly, first learning tasks with low capability demands and then gradually increasing complexity. This method is named Routing-Guided Curriculum. Training also adds On-Policy Distillation: first the student solves problems in its own way, then the teacher provides guidance targeting the actual steps the student took. Results show that Agent