
Chinese technology company Xiaomi announced that it has opened invitation-only testing for MiMo Desktop, a desktop AI application that executes complete workflows rather than generating responses in a chat interface. The release positions the company as the first China-based laboratory to offer computer-use capabilities to users through its flagship models.
Real-world work rarely begins with a well-defined prompt. It typically involves spreadsheets, images, video, PDF documents, audio recordings, and compressed files that collectively constitute the context of a task. MiMo Desktop accepts these multi-format inputs directly, without requiring prior organization or format conversion. Users state their objectives in natural language, after which the system interprets the materials, decomposes the task, invokes appropriate tools, and delivers editable outputs — including documents, spreadsheets, presentations, web pages, audio, video, 3D models, and software engineering projects.
Two design decisions distinguish the product. First, the result preview is not a static rendering but a fully interactive deliverable incorporating component structure, interaction logic, and data visualization, which users can operate and revise within the session — whether the output is a data dashboard, a game prototype, or a presentation. Second, modifications are made by selection: users highlight a chart, paragraph, or region and describe the desired change, and only that portion is regenerated. Every revision is preserved in version history, allowing users to compare drafts and roll back to previous states.
The system also incorporates “Smart” scheduling. Rather than requiring users to select a model manually, MiMo Desktop assesses task type, complexity, and cost requirements, routing routine work to standard models for speed and complex, multi-step tasks to flagship models for higher output quality. Large-scale tasks can be distributed across multiple agent sessions that collaborate while maintaining independent memory and workspaces. Xiaomi reports that cache optimization achieves hit rates of up to 99% within a session on long tasks, which reduces redundant computation and controls operational costs.
The second dimension of the release concerns control. MiMo Desktop operates the browser as both an information source and an execution environment: it opens pages, retrieves information, completes forms, extracts assets, and imports findings directly into the task. When producing web-based outputs, the system can additionally verify deliverables through the browser during the workflow.
The most important feature is full computer control, available in the overseas version. MiMo Desktop reads screen content and operates the keyboard and mouse across applications — opening files, verifying data, and transferring information between programs — with built-in result verification that can trigger corrections or halt execution when necessary. For stable and repetitive processes, a Record & Replay function allows users to demonstrate a workflow once and have the system re-execute it through natural language instructions.
This reflects a deliberate architectural approach: the desktop is treated as the operating environment, and the model functions as an agent within it, rather than as a language interface layered on top of existing applications.
Computer-use capability is not primarily an interface problem but a technical one, and Xiaomi has been developing the underlying foundation for more than a year. The groundwork was established in mid-2025 with MiMo-VL-7B, the company’s open-source vision-language model, which achieved a then-record score of 56.1 on OSWorld-G — the standard benchmark for GUI grounding — outperforming models developed specifically for the task, such as UI-TARS. This capability was developed through a four-stage training pipeline covering mobile, web, and desktop interfaces, together with a substantial corpus of Chinese GUI data and a training task that infers intermediate actions from before-and-after screenshots — precisely the perceptual foundation required for a computer-use agent.
The model lineage has continued to develop. By early 2026, Xiaomi’s flagship MiMo-V2-Pro ranked in the same tier as Claude 4.5 Sonnet, GPT-5.2, and Gemini 3.0 Pro on coding agent, general agent, and tool-use benchmarks. This provides the Desktop client with a top-tier model for complex tasks while the routing layer maintains cost efficiency for routine operations. In addition, Xiaomi brings more than two decades of experience in consumer hardware and software ecosystems, which is reflected in the product’s focus on the practical conditions of a user’s desktop environment rather than on benchmark performance alone.
For observers of the AI sector, the release indicates a broader shift in evaluation criteria: the frontier is moving from reasoning benchmarks toward operational competence — agents that complete forms, verify their own output, and remain within defined compute budgets. Xiaomi’s combination of GUI-grounding models, a frontier-tier flagship, and a computer-use interface now in beta places the company among the leading contenders in the emerging category of autonomous desktop agents.
The post Xiaomi Becomes First China-Based Lab To Offer Computer Use, Unveils MiMo Desktop Agent In Invite-Only Beta appeared first on Metaverse Post.