---
postId: research.002
lang: en
---

# Manifold and the Overseas World-Model Landscape

> A post-project account of how three Deep Research inputs became an auditable, falsifiable, and reproducible decision chain. The research snapshot is fixed at **July 14, 2026**; it covers four non-China startups and one large-company project.

This is an English companion to the original Chinese report. It preserves the same decisions, evidence boundaries, and audit route in a shorter form. The complete Chinese narrative and the full sanitized project snapshot remain the canonical record.

The five final research objects were **World Labs, Odyssey, Runway, Decart, and NVIDIA Cosmos 3**.

> **Evidence boundary.** This work is an independent desk-research archive. It does not represent Manifold AI or any company discussed here, contains no company-provided internal material, and does not turn vendor or partner statements into proof of paid production adoption.

---

<a id="en-quick-take" data-pair-id="quick-take"></a>
## One-minute takeaway

1. **Supply has split into three observable paths:** publicly priced interfaces, application-controlled access, and downloadable or self-hosted components.
2. **The final five are not a ranking.** World Labs, Odyssey, Runway, Decart, and NVIDIA Cosmos 3 represent five different product and platform risk paths.
3. **The evidence ceiling remains low.** As of July 14, 2026, the research found no customer-first-party confirmation of paid production adoption, renewal, or revenue for any of the five. Callable is not the same as scaled deployment.
4. **Manifold is only a public comparison anchor.** Without internal customer, revenue, standard-delivery, or roadmap evidence, strategic relationships remain conditional rather than claims of active competition for the same buyers or budgets.
5. **The practical next step is validation and rights infrastructure:** define standard delivery units; build shared tests for action fidelity, persistent state, geometry and collision, failure rate, latency, and TCO; and establish data, model-improvement, deployment, and redistribution rights in advance.

> **Evidence model.** The archive separates source facts, company reports, partner-first-party disclosures, independent counter-evidence, conditional analysis, and unresolved questions. No source is allowed to support more than its actual atomic claim.

<a id="en-design-ownership" data-pair-id="design-ownership"></a>
## 00 · Research design: turn a broad question into a verifiable decision system

I did not hand one broad prompt to an AI and wait for a market report. I defined the question, split it into three Deep Research tasks, set geographic and technical qualification gates, designed approval checkpoints and evidence structures, and retained responsibility for red-team revisions, the final five-object selection, and the public wording. AI and agents were the execution layer for parallel research, source recovery, structured organization, and mechanical QA.

The core design separated four states that are often collapsed: **worth investigating → technically qualified → strategically related → worth one of five presentation slots**. Manifold was used only as a public comparison anchor because public information could not establish its actual customer set, revenue, standard delivery unit, or internal roadmap.

Every visible conclusion had to pass through a `Source Card → Atomic Claim → Evidence Map` chain. Checkpoints, red-team review, and content lock kept the speed of AI execution subordinate to explicit human approval.

<a id="en-research-origin" data-pair-id="research-origin"></a>
## 01 · Why the project began with three separate Deep Research tasks

The initial work was split into three mutually checking questions:

- **DR-01 — describe Manifold from public evidence.** It established dimensions for comparison but could not establish private commercial facts, so Manifold remained an anchor rather than a competition node.
- **DR-02 — broaden the non-China discovery pool.** It supplied 23 object leads, but the word “world model” covered incompatible meanings across video, robotics, simulation, and spatial intelligence. Discovery therefore could not equal qualification.
- **DR-03 — map large-company projects and substitution routes.** It prevented a startup-only view while keeping VLA systems, simulation platforms, and internal components from automatically occupying a world-model slot.

The project plan and launch prompt defined hard boundaries, approvals, and evidence rules. All three reports were treated as leads; consequential claims had to be reopened at their original public sources.

<a id="en-decision-trail" data-pair-id="decision-trail"></a>
## 02 · The complete decision trail

The public decision chain was:

1. **Raw leads:** three Deep Research reports with broad but uneven coverage.
2. **Input audit:** separate factual leads, analytical judgments, and sources that needed recovery.
3. **Object cleaning:** apply geography, object type, technical qualification, and component-isolation rules.
4. **Qualification and selection:** move from 23 leads to 10 validation objects, then to five focus objects, four watch objects, and one early signal.
5. **Five dossiers:** search for supporting evidence, counter-evidence, component boundaries, and unresolved questions for every selected object.
6. **Red team and content lock:** attack version ownership, access, licensing, adoption language, and scope; then map each page claim back to atomic claims and sources.
7. **Research repository:** preserve inputs, decisions, evidence, QA, and generation artifacts so the work can be reproduced without relying on chat history.

The point was not to make the process look elaborate. Each stage removed a different failure mode before the next stage was allowed to begin.

<a id="en-governance-first" data-pair-id="governance-first"></a>
## 03 · Step one: establish authority and state before searching

The initial reports already contained many conclusions. Without a source-of-truth hierarchy, later work would have mixed raw input, newly recovered evidence, analysis, and approved decisions.

The project therefore began with a machine-readable state file, decision log, and input manifest. Inputs were recorded with identity, size, date, and SHA-256. A single canonical editor owned authoritative files, while scripts handled hashes and mechanical checks. AI was useful here because it translated a natural-language charter into explicit states, enumerated boundaries, and stop conditions—not because it wrote an early conclusion.

<a id="en-split-audit" data-pair-id="split-audit"></a>
## 04 · Step two: audit the three research inputs separately

Three bounded audit agents examined the Manifold anchor, the non-China universe, and large-company alternatives. They returned conflicts and recommendations; the main review reconciled them under one standard.

The 23 DR-02 entries became **nine original validation inputs, thirteen boundary exclusions, and one discarded item that could not be restored to a qualified object**. This was a key result: most of the early work was not “finding 23 companies,” but removing category contamination before deep-research resources were spent.

<a id="en-qualification" data-pair-id="qualification"></a>
## 05 · Step three: define a qualifying world-model object and reopen original sources

Vendor use of the phrase “world model” was insufficient. The technical gate required action, control, or persistent interaction to have an auditable, temporally consistent effect on the represented future state. Navigation-qualified state could include agent or camera pose, reachability, collision constraints, and persistent memory; the rule could not require only object manipulation.

The strategic gate was separate. A technically qualified object entered the core set only if public evidence supported a product, platform, or credible future-market relationship to the Manifold public anchor. Search agents reopened official project pages, papers, repositories, model cards, pricing, terms, licenses, and partner statements rather than citing the Deep Research paraphrases.

This stage added an omitted object, refreshed version ownership such as V-JEPA 2.1, isolated Manifold’s auxiliary assets from unproven commercial priority, and kept insufficiently evidenced objects in an isolation or early-signal layer.

<a id="en-portfolio-selection" data-pair-id="portfolio-selection"></a>
## 06 · Step four: validate each object and select five without ranking them

Technical qualification, external availability, and commercial competition were tested independently. For each object, the research asked whether actions changed the modeled future, whether the entity met the geographic rule, how an outsider could obtain the system, who reported the adoption evidence, and whether strategic relevance remained conditional.

The final portfolio was deliberately complementary rather than ranked:

- **World Labs:** spatial-world products and API plus real-time navigation research;
- **Odyssey:** a real-time world-stream developer service;
- **Runway:** a content platform branching into Worlds and Robotics;
- **Decart:** a publicly priced, driving-first preview API;
- **NVIDIA Cosmos 3:** downloadable and self-hostable action-conditioned components.

Runway retained a slot over Genie 3 because it supplied a public application path and product surface, while Genie 3 remained a frontier capability and internal-platform watch item. Cosmos 3 represented the self-hosted platform route, preserving four startup slots without ignoring large-company substitution risk.

<a id="en-dossier" data-pair-id="dossier"></a>
## 07 · Step five: use dossiers to test what had not yet been proven

The five dossiers shared four questions: what first-party evidence established, what remained vendor- or partner-reported, which components could not borrow one another’s capabilities or access status, and what future evidence would change the judgment.

The research separated RTFM from Marble and the World API, Odyssey-2 Pro from Max and other previews, Runway Worlds from Robotics, and Cosmos models from runtimes and product containers. A partner test could not silently become proof of paid production adoption.

This phase produced five reader-facing dossiers, **58 targeted support and counter-evidence queries, 141 registered sources and source cards, and 111 atomic claims**. The most consequential negative finding was consistent across all five: no object had been restored to customer-first-party confirmation of paid production deployment, renewal, or revenue as of the cutoff.

<a id="en-red-team" data-pair-id="red-team"></a>
## 08 · Step six: compare without false precision, then red-team the result

The five objects did not share hardware, task, duration, failure-rate, or commercial-disclosure protocols. A total score would therefore have manufactured precision. The comparison stayed on four auditable dimensions: technical evidence, external availability, commercial-adoption ceiling, and conditional relevance to Manifold.

Independent reviews attacked technical versioning, commercial access and licensing, and scope and narrative boundaries. This downgraded the strategic gate to an explicitly conditional public-anchor comparison, corrected overstatements about pilots and hands-on evidence, and generated **17 monitoring signals for the final five plus five triggers for watch objects**.

<a id="en-content-lock" data-pair-id="content-lock"></a>
## 09 · Step seven: content lock made the report traceable

Readable prose can still drift if its visible claims cannot be mapped to evidence. The project therefore created a `Page Claim → Atomic Claim → Source` map, kept CSV as the authoritative representation, generated a review workbook, and froze the relevant Markdown, CSV, XLSX, and ledger files with hashes.

This process caught concrete errors: an NVIDIA NIM version attribution, three different Odyssey time limits that had been conflated, aggregated evidence where object-level mapping was required, and cross-page borrowing of license, API, action-mode, or component claims. The locked result contained **63 unique page-evidence mappings across ten content units**.

<a id="en-deck" data-pair-id="deck"></a>
## 10 · Step eight: translate locked findings into a deck

After content lock, the project converted the evidence into a ten-page storyboard and PowerPoint. Structured generation handled repeatable layout; Microsoft PowerPoint rendered at 1280×720; programs checked text and page bounds; and human review examined hierarchy, wrapping, and visual emphasis at full size and as thumbnails.

Four PPTX versions and several visual-correction passes produced a final ten-page deck whose 252 text boxes passed mechanical and human QA. The deck compresses the conclusions; the archive preserves how those conclusions were reached.

<a id="en-repository" data-pair-id="repository"></a>
## 11 · Step nine: convert a one-off project into a durable research repository

The closing phase separated original inputs, the sanitized historical workspace, reader-facing method and result documents, agent and tool provenance, and reproducibility scripts. Of 333 historical workspace files, 331 public-safe files were retained; two derived ZIP byte packages with local-machine metadata and no unique content were excluded while their hashes and sixteen member files remained available.

The migrated portfolio entry now preserves the source snapshot, including original inputs, dossiers, source cards, ledgers, decision records, red-team work, deck generations, and QA. The source commit and migration checks are recorded separately so future deletion of the standalone repository does not remove the substance of the work.

<a id="en-postmortem" data-pair-id="postmortem"></a>
## 12 · What the project taught me

The most useful choices were separating the three opening questions, treating discovery, qualification, strategic relation, and presentation selection as different states, recording support, counter-evidence, and unknowns in the same system, and putting checkpoints and content lock between AI execution and public claims.

If repeated, I would design the delivery and retrospective layers together on day zero, encode every approval as executable state, attribute each claim to its discovering agent earlier, pressure-test the densest Chinese typography before building the full deck, and measure AI-native value by risks closed and omissions prevented—not by agent count.

<a id="en-ai-value" data-pair-id="ai-value"></a>
## 13 · What AI actually added

AI contributed in four defensible ways: parallel coverage across objects and evidence types; structured memory in cards, ledgers, registers, and decision logs; adversarial review that searched for version drift and semantic escalation; and mechanical validation of identities, hashes, workbooks, evidence maps, and rendered bounds.

It did not replace scope ownership, stage approval, the five-object selection, content lock, or the judgment required to phrase claims at their evidence ceiling. The session recorded 33 agent-creation requests and 30 actual launches, but the number matters only when each agent solved a bounded problem and left a verifiable artifact.

<a id="en-conclusions" data-pair-id="conclusions"></a>
## 14 · Final conclusions

Overseas world models had moved beyond demonstrations into three observable delivery paths: **publicly priced interfaces, application-controlled access, and downloadable or self-hosted components**. The competitive unit had expanded from visual quality to interfaces, runtimes, deployment, integration, data rights, licensing, and total cost of ownership.

Yet the public record at the research cutoff did not establish customer-first-party confirmation of paid production adoption for any of the five selected objects. A safer action set for Manifold was therefore to:

1. define standard delivery units across model, API, deployment service, and joint project;
2. build shared validation assets for action fidelity, persistent state, geometry and collision, failure rate, latency, and TCO;
3. establish input/output data rights, model-improvement rights, deployment, and redistribution rules before commercial scale obscures them.

<a id="en-audit" data-pair-id="audit"></a>
## 15 · How to audit the research

Follow the actual decision chain rather than browsing by file count:

1. read the three Deep Research inputs;
2. compare the input manifest and eligibility register;
3. sample the source registry, source cards, and evidence ledger;
4. read the five dossiers and cross-comparison;
5. inspect the red-team report for claims that were downgraded, corrected, or retained;
6. run the repository audit script against the migrated snapshot.

The ten-page final deck is available at [`source-snapshot/archive/manifold-world-model-landscape/10_deck/manifold_world_model_landscape_v1_4.pptx`](source-snapshot/archive/manifold-world-model-landscape/10_deck/manifold_world_model_landscape_v1_4.pptx), and the original source snapshot begins with [`source-snapshot/README.md`](source-snapshot/README.md).

<a id="en-current-boundaries" data-pair-id="current-boundaries"></a>
## Appendix · Current boundaries

- This is a public research-process record fixed at the stated cutoff, not a live market database.
- Original web pages can change; source cards preserve the support and limits observed during the research.
- The raw Codex conversation was not published because it included system data and unrelated later work; sanitized agent, tool, decision, and evidence records were retained instead.
- No PDF was produced in the original project. PPTX v1.4 is the final ten-page conclusion artifact.

Core first-party entry points retained in the source registry:

- [Manifold WorldScape](https://manifoldai.cn/blogs/WorldScape.html)
- [World Labs](https://www.worldlabs.ai/about)
- [Odyssey](https://odyssey.ml/the-gpt-2-moment-for-world-models)
- [Runway GWM-1](https://runwayml.com/research/introducing-runway-gwm-1)
- [Decart Oasis](https://decart.ai/oasis)
- [NVIDIA Cosmos 3](https://research.nvidia.com/labs/cosmos-lab/cosmos3/)
