We believe the release of Xiaomi MiMo presents an opportune moment to articulate its positioning within the existing large model ecosystem and analyze its differentiating characteristics compared to other mainstream large models.
Our core vision is to build a system that not only possesses the capability to process massive amounts of information, but can also achieve alignment with human intent at a deeper level. We are committed to bridging the gap between large models and human cognition, enabling models to not only execute instructions efficiently, but also accurately understand contextual nuances and users' emotional psychology, thereby becoming more beneficial and reliable collaborative tools.
Therefore, we are conducting this special assessment. Although the reference value of standardized benchmark scores is widely recognized, we still believe that evaluation dimensions should extend to areas that are more difficult to quantify yet crucial to human experience—namely aesthetic judgment and emotional resonance capabilities.
In constructing the evaluation framework, we deliberately selected two models (Model-1 and Model-2 henceforth) representative of the current industry ecosystem as comparisons, aiming to establish a high-standard comparative benchmark to more precisely anchor MiMo's performance and differentiating characteristics. To effectively assess these qualities, 9 evaluators utilized their respective domain expertise, intuition, and judgment criteria to review MiMo and the two comparison models, ultimately producing the content in this assessment report.
In the realm of creative writing, we analyzed MiMo's performance across various tasks and the underlying capabilities it employed, characterizing its overall creative style.
To evaluate MiMo against the two comparison models, we designed writing tasks across three distinct literary genres: modern poetry, classical Chinese poetry, and narrative fiction. The outputs were then analyzed against our framework of core creative writing dimensions—including imagination & originality, logic & structure, emotional resonance, linguistic finesse, knowledge application, and instruction following—to assess MiMo's overall performance and capability distribution.
The evaluation comprised a total of 40 tasks: 15 each for modern poetry and classical Chinese poetry (all in Chinese), and 10 for narrative fiction (3 of which in English). For each task, outputs from the models were manually graded on a 1-to-5 integer scale. The table below outlines the key criteria for each genre across the six core dimensions.
The images below detail the performance of MiMo against the two comparison models. Overall, MiMo stands out as a relatively stable and balanced model. While Model-1 leans heavily on logic and Model-2 on emotion, MiMo skillfully integrates both logical structure and emotional depth. Furthermore, it demonstrates solid instruction-following capabilities, vivid imagination, and excellent linguistic finesse.
MiMo's creative capabilities adapt flexibly to different literary genres. For modern poetry, it places a greater emphasis on imagination and emotional resonance. In contrast, for narrative fiction, it strengthens its focus on logical coherence and structural integrity. However, MiMo's performance in classical Chinese poetry reveals an area for further development, as manifested in its inconsistent adherence to prosodic rules and its struggle to reliably apply relevant knowledge.
Our observations so far portray MiMo as a model that is both balanced and adaptable in literary creation. It is anchored in strong capabilities for empathy and emotional expression, and complemented by logic and imagination. In doing so, it seeks to adopt the most fitting authorial voice for any given genre.
In the domain of artistic perception, we focus on MiMo's specific performance across different types of aesthetic and creative tasks.
Assessment tasks were constructed around three core modules: aesthetic preferences, art knowledge discussion, and painting/design creation, evaluating MiMo and two comparison models. The assessment covers diverse artistic scenarios, with creative questions accounting for over 50%, including painting, illustration, graphic design, spatial and installation art, with emphasis on testing models’ creative generation and visual description capabilities. Additionally, this assessment integrates cross-cultural art history theory and anthropomorphized aesthetic communication, accommodating both Chinese and English questions.
Based on the decomposition of core dimensions of comprehensive artistic capability, evaluation conclusions were mapped to different dimensions to further observe and analyze the models' overall performance and capability distribution.
The specific performance of MiMo and the two comparison models is shown above. In artistic literacy, MiMo demonstrates comprehensive balancing capability and systematic thinking advantages.
Compared to Model-1's focus on rational high execution but somewhat conservative approach, and Model-2's focus on high anthropomorphism and divergent creativity but difficulty in practical implementation, MiMo found a balance between rational analysis and emotional expression—maintaining objectivity in professional analysis while retaining warmth in emotional connection, along with excellent logical architecture and evidence integration capabilities.
Based on current observations, we believe MiMo is a model that achieves organic integration of logic and avant-garde innovation in artistic creation. It uses robust systematic thinking as its framework, supplemented by certain visual transformation capabilities. Although there is still room for deeper exploration, it successfully avoids mediocrity or impracticality in solutions, demonstrating significant development potential.
In the field of philosophy, we focus on MiMo's ability to guide interactions for non-professional users, examining whether it can effectively inspire users to engage in deep thinking.
This evaluation constructs tasks based on simulating philosophical inquiry from a non-professional perspective.The dataset consists of 35 questions, spanning both Chinese and English, covering topics across AI philosophy, ethics, metaphysics, and epistemology. It encompasses both cutting-edge discussions on AI ethics and machine cognition, as well as classic topics such as free will, personal identity, skepticism, and the ontology of fictional characters. Based on the decomposition of core dimensions for comprehensive philosophical interaction capabilities, we mapped evaluation conclusions to different dimensions according to four progressively deeper assessment criteria: content readability, factual accuracy, key point coverage, and depth of insight, to further observe and analyze the models’ overall performance and capability distribution.
The specific performance of MiMo and comparison models is shown in the figure above. Overall, although all three provided high-quality answers to the questions, their underlying 'thinking personalities' were distinct: MiMo is a patient and guiding philosophy popularizer, Model-1 is a serious and aloof scholar, while Model-2 tends toward emotional expression, like an eloquent orator.
Specifically, at the content readability level, we examined the models' ability to establish a conversational feel and popularize professional concepts. MiMo performed excellently in this regard, not only with a lively style but also demonstrating excellent guidance; in comparison, Model-1, while rigorous, appeared somewhat stiff and obscure with poor interactivity, while Model-2, though warm, was slightly verbose. Secondly, at the factual accuracy level, we strictly verified the credibility of philosophical assertions and scientific citations provided by the models. Results showed that all models occasionally had flaws, with Model-2 having the most significant factual errors. Finally, regarding key point coverage and depth of insight, we focused on evaluating the models' precision in capturing user confusion, the completeness of theoretical frameworks, and philosophical insights. The evaluation found that the key to high-quality responses lies in precisely grasping core issues and uncovering profound philosophical insights. Among them, Model-1 had the best theoretical reserves and performed better than MiMo and Model-2 under stable conditions. However, in a few cases, it failed to clearly identify user intent, resulting in off-topic responses. In comparison, MiMo's performance was more robust, achieving high-quality output while ensuring relevance.
Based on current observations, we view MiMo as a model that strikes an organic balance between "popular readability" and "theoretical rigor" in philosophical dialogue. Although there is still room for improvement in the sharpness and depth of viewpoints, its stable performance makes it an excellent choice for guiding non-professional users in philosophical thinking.
In the field of screenplay writing, this evaluation focuses on aesthetic perception and execution efficiency in film and television scriptwriting, examining creative performance across multiple domains.
Unlike single-dimension text generation tests, we focus on the model's comprehensive performance in full-chain creative scenarios, with examination scope spanning from high-EQ interactions involving logical traps and social nuances, to structural stability in long-form narratives exceeding 5,000 words. We pay special attention to whether the model can maintain setting coherence and reasonable empathy across the large span from short story creativity to long-text engineering, as well as the precision of implementation when facing specific style instructions, to determine MiMo's capability characteristics in scriptwriting scenarios.
This evaluation constructed 37 targeted tasks across three core screenwriting dimensions: Setting Originality (12 questions), Textual Richness (13 questions), and Instruction Control (12 questions), with clearly established circuit-breaker and bonus mechanisms.
For each question, responses from MiMo and two comparison models were manually scored on an integer scale of 1-5. For Setting Originality, we not only identified creative generation but also assessed whether creativity closely empowers the story itself, requiring insight into the subtleties of human nature even in extreme situations. Textual Richness aims to quantify narrative granularity, focusing on whether the model can discern the hidden purpose of questions and elevate the overall atmosphere through precise style matching and memorable lines. Instruction Control examines logic adherence and style transfer in depth while establishing baselines for physical laws and historical facts, requiring models to respond flexibly to complex logical traps while avoiding common-sense errors and temporal-spatial inconsistencies.
MiMo and Model-1 achieved comparable overall scores, but exhibited distinctly different capability emphases and limitations. MiMo ranked first in both Originality (C1) and Textual Richness (C2); however, its weak logical chains (C3) often led to poor handling of hard constraints in questions due to excessive pursuit of creative uniqueness. In contrast, Model-1 performed better in narrative construction and style imitation, with stronger capability for producing memorable lines, but its low creativity score (C1) revealed generally weak empathy, making its creations relatively lacking in warmth.
Model-2, in the trailing position, maintained logical closure in basic dialogue but clearly fell behind in advanced creative dimensions. Its lack of creativity (C1) made it difficult to explore human depth, with text easily falling into verbose and formulaic happy-ending patterns, limiting its application potential in professional screenplay development.
Based on current observations, MiMo resembles an experiential model that prioritizes aesthetics over constraints in text creation, attempting to construct a creative persona with cinematic quality and dramatic tension through higher originality and relatively nuanced empathic perception.
This test aims to examine LLMs’ critical thinking capabilities regarding social phenomena. The test includes 18 questions on interpretation (explaining social mechanisms) and 12 questions on policy and instruction(what governments, enterprises, and individuals should do). Among these, 24 are in Chinese, with 6 in English, German, and French (addressing international issues). This evaluation focuses on models’ critical thinking capabilities across five dimensions: 1. If they can accurately understand the problems; 2. If there are definitions of concepts and related theories; 3. Whether arguments cover various perspectives; 4. Accuracy and validity of arguments; 5. Reflection on the Q&A itself. Additionally, styles and characteristics of each model are discussed as well.
[Interpretation] If a city provides higher assistance payments to unemployed individuals, some worry this will reduce motivation to find work, while others say it gives people more confidence to transition careers. What's your view?
[Instruction] As a company using AI to screen resumes, what policies would you set so hiring stays efficient but applicants still feel the process is fair and explainable?
According to Max Weber, the founding figure of sociology, understanding social phenomena requires grasping the subjective meaning that humans assign to their behaviors, while avoiding researchers' mixing their own subjective values into objective analysis. Examining LLMs’ critical thinking capabilities involves examining whether they can perceive the multilayered structures of human society and read social meanings from different standing points.
Since social observations have no definite answers, this test adopted a relative way of scoring. Tester first provides a medium-performance response sample, then compares model responses against this sample to determine whether performance in each dimension was “better” or “worse,” and finally assigns scores. Results are shown in the figure below.
Overall, Model-1 demonstrated comprehensive critical thinking capabilities, performing best in problem understanding and argument validity, with impressive performance on interpretative questions; however, it showed weaker reflexivity, tending to make overgeneralized judgments or lacking reflection on its own standing point. Model-2 performed relatively well on reflective criticism dimensions but lacked organization between viewpoints and showed slightly weaker performance on international issues. MiMo frequently responded by listing items or enumerating pros and cons, lacking critical examination of viewpoints on interpretation questions, performing better on recommendation questions, and showing stable performance on international issues.
All three models performed relatively weakly on concept definition and reflective criticism dimensions. This suggests we should focus on improving model performance in these directions, emphasizing reflection on the questions themselves and on one’s own standpoint.
This test aims to simulate scenarios where users utilize large language models for legal studies, thereby exploring the legal knowledge reserves and reasoning capabilities of models like MiMo in such contexts. This test constructed 40 questions across four sections: legal maxim comprehension (10 questions), knowledge Q&A (15 questions), case analysis (5 questions), and simulated decision-making (5 questions). After obtaining responses, manual scoring was conducted across five dimensions: knowledge accuracy, logical reasonableness, argumentation appropriateness, linguistic comprehensibility, and viewpoint clarity. The evaluation results are shown in the figure below.
Overall, in the legal domain, MiMo is a model that leans toward knowledge reserves and appropriate argumentation. Compared to Model-1's frequent insufficient reasoning and argumentation and Model-2's emphasis on language expression with occasional knowledge gaps, MiMo relatively well balances both knowledge breadth and argumentation depth, while also possessing excellent professional language command and relatively robust legal viewpoint output capabilities.
On legal maxim comprehension questions, MiMo demonstrated strong understanding ability, but its output stability was inferior to Model-2; knowledge Q&A questions are MiMo's strong suit, where it demonstrates proficient mastery of legal terminology and knowledge points and can elaborate with vivid language and rigorous argumentation frameworks, though its logical performance was not as good as Model-2; in case analysis questions, despite facing high-difficulty problems where MiMo's logic scores were poor, its argumentation scores enabled it to handle analytical argumentation development with relative proficiency; on simulated decision-making questions, MiMo was able to fully and appropriately elaborate on its decision rationale, demonstrating good instruction-following ability.
It can be said that when studying law, MiMo will assist you with its solid knowledge reserves and appropriate exposition methods, and will timely use vivid examples to help you understand difficult concepts. It can effectively grasp contextual needs and switch between accessible and professionally profound writing styles. However, when facing difficult cases in case analysis, its logical performance is relatively poor, making it difficult to reach correct conclusions amid complex legal relationships.
This evaluation is built upon our ongoing research into balancing safety and user experience in large language models. The research aims to explore how models can provide maximum effective feedback to users while ensuring safety. We focus on model performance in identifying and rejecting harmful requests, while also assessing whether they can reasonably satisfy user needs within safety boundaries and minimize unnecessary obstacles that safety mechanisms may pose to user experience.
In this evaluation version, we included MiMo and two other mainstream models. The test covered 500 safety test scenarios, with all models evaluated on the same test set, generating a total of 14,340 fine-grained judgment records. The test design covered 190 different types of safety threats, primarily distributed across 8 safety categories: sexual content, self-harm and suicide, harassment and hate, controlled substances, criminal planning, fraud and deception, and unauthorized advice.
It should be noted that AI safety evaluation itself is a highly complex and continuously evolving field. Although this test covered the main currently known risk types, it is still difficult to exhaust all potential safety scenarios and attack variants. Therefore, the evaluation design in this article is primarily intended to provide a fair and reproducible basis for relative comparison of different models under the same conditions. For more complete evaluation assumptions, safety boundary definitions, and methodological details, please visit our detailed safety blog to review the complete AI Safety documentation.
At the end of this evaluation, we invited 9 evaluators to describe their impressions of MiMo, resulting in the following impression word cloud.
In this evaluation, MiMo not only demonstrated capabilities comparable to existing mainstream large models, but also developed differentiated characteristics of flexibility in general domains due to its outstanding empathy capabilities compared to other models.
If you wish to learn about complete model details, please read our technical report.
At 1001, we have more excellent Q&A cases from MiMo.
We look forward to you interacting with MiMo directly through the website, or calling the API to get responses from MiMo.
Limitations Statement: The test set construction and manual scoring for this evaluation were completed by 9 evaluators with relevant disciplinary expertise. Given the relatively small sample size, the evaluation results may contain certain biases. Additionally, constrained by specific disciplinary paradigms, the evaluation criteria inevitably carry cognitive tendencies and subjective emphases inherent to these professional fields. Users are advised to fully consider these contextual factors when interpreting the data, and to examine the evaluation results within their specific professional contexts, thereby establishing more objective and reasonable expectations regarding model capabilities.
Welcome to join us if you are interested in our project. See more on the JD page of Knowledge Engineer.