
What is LingBot-Depth 2.0?
LingBot-Depth 2.0 is a next-generation foundational model for robotic spatial perception released on July 7, 2026, by Robbyant, a subsidiary of Ant Group. It is not positioned as a large language model, but rather as a foundation model for 3D vision, primarily designed to enable robots to ”see the three-dimensional world clearly.”
GPT is the brain of AI, while LingBot-Depth is the robot's eyes.It addresses one of the most fundamental issues in robotics—Depth Perception.
Technical Principles of LingBot-Depth 2.0
The breakthroughs achieved by LingBot-Depth 2.0 are primarily attributable to the following two core technologies:
- Spatial Native Vision Platform (LingBot-Vision): It is the industry’s first visual foundation model to use “boundary structure” as a pre-training objective. By introducing geometric constraints and a boundary-centered masking mechanism, it possesses sub-pixel-level boundary localization and spatial structure understanding capabilities, providing robust foundational support for depth estimation.
- Masked Deep Modeling (MDM) and Big Data: The model uses the “depth gaps” naturally created by depth sensors on transparent and reflective objects as masks, requiring the model to “fill in” the complete depth using only RGB color images and the remaining valid depth information. At the same time, the scale of its training data has expanded significantly from 3 million in version 1.0 to 150 million, creating a virtuous cycle between data and the model.
Technological Advantage:
| abilities | LingBot-Depth 2.0 |
|---|---|
| RGB + Depth Fusion | ✅ |
| Metric Depth | ✅ |
| Advanced Autocomplete | ⭐⭐⭐⭐⭐ |
| Transparent Objects | ⭐⭐⭐⭐⭐ |
| Mirror | ⭐⭐⭐⭐⭐ |
| Small Object Recognition | ⭐⭐⭐⭐☆ |
| Edge Sharpness | ⭐⭐⭐⭐⭐ |
| Robotic Gripping | ⭐⭐⭐⭐⭐ |
| 3D reconstruction | ⭐⭐⭐⭐⭐ |
| SLAM | ⭐⭐⭐⭐☆ |
Key Features of LingBot-Depth 2.0
- In-Depth Autocomplete for Complex Materials: It effectively overcomes the “hard scenes”—such as glass, mirrors, and transparent objects—where traditional depth cameras often fail, and is capable of generating 3D depth maps with complete structures and clear boundaries.
- Precise Spatial Awareness: It delivers outstanding performance in edge clarity, recognition of small objects (such as cables and thin rods), and long-range depth estimation, thereby preventing obstacle-avoidance blind spots caused by the robot missing objects.
- High Precision and High Robustness: In the most challenging indoor scenarios involving large-scale depth loss, the depth error (RMSE) was halved compared to the previous generation (from 0.132 to 0.062), and the system ranked first in 12 out of 16 depth completion benchmark evaluations.
Use Cases for LingBot-Depth 2.0
- Embodied Intelligence and Robotics: Provides a reliable spatial data foundation for robot navigation, obstacle avoidance, and precise grasping (such as grasping transparent storage bins, glass cups, and champagne towers).
- industrial automation: Enables tasks such as precision assembly, alignment of highly reflective metal parts, and tracking the edges of cardboard boxes in logistics and warehousing.
- AR/VR and Autonomous Driving: Provides precise depth information to enhance safety and user experience in environments with complex lighting and materials.
LingBot-Depth 2.0 project repository
- Project website::https://technology.robbyant.com/lingbot-vision
- GitHub repository::https://github.com/robbyant/lingbot-vision
- HuggingFace Model Library::https://huggingface.co/collections/robbyant/lingbot-vision
- arXiv Technical Paper::https://github.com/robbyant/lingbot-vision/blob/main/paper.pdf
How do I use LingBot-Depth 2.0?
- Get the source code: The model weights for LingBot-Vision (available in four versions: ViT-G, L, B, and S) and its technical report have been fully open-sourced; developers can access them via GitHub, Hugging Face, or ModelScope.
- Hardware and SDK Integration: Ant Lingbo has entered into a deep partnership with Orbbec. Users can enable spatial perception capabilities with a single click using Orbbec’s SDK, or plan to purchase an all-in-one “3D camera + spatial perception” camera product—which integrates the commercial version of this model—by the end of the year for out-of-the-box use.
- Cloud Deployment: Developers can also leverage AI computing platforms such as Baidu Baige to quickly run models and perform inference using the official Docker images provided.
Comparison of similar products
| offerings | firms | localization | dominance | inferior |
|---|---|---|---|---|
| LingBot-Depth 2.0 | Ant Lingbo | Spatial Perception Foundation Model | Outstanding performance with transparent objects, mirrors, and depth restoration | The ecosystem is still under development |
| Intel RealSense SDK | Intel | Deep Learning Algorithms | Mature and stable | The glass performed averagely |
| Orbbec SDK | Aobi Zhongguang | Camera Algorithms | Domestic Ecosystem | More reliance on hardware |
| NVIDIA Isaac Perceptor | NVIDIA | Robotic Vision | Seamless integration with the Isaac platform | Heavily reliant on the NVIDIA ecosystem |
| Meta SAM2 | Meta | Image Segmentation | Strong segmentation capabilities | Does not provide genuine depth |
| Depth Anything V2 | Depth Estimation | Monocular Depth | No depth camera required | Not Actual Scale (Metric Depth) |
| Metric3D | Tsinghua University, etc. | Monocular Metric Depth | High accuracy | Limited support for transparent objects |
| FoundationStereo | Meta, etc. | Stereoscopic Depth | Both eyes are estimated to be excellent | Complex materials still pose a challenge |
The biggest difference is:
LingBot-Depth is not a monocular depth estimation model, but rather a ”depth completion + depth enhancement” model designed for RGB-D cameras, with a greater emphasis on spatial perception in scenes featuring real-world scales and complex materials.
data statistics
Relevant Navigation

Tencent introduced the industry's first open source world model that supports native 3D reconstruction and ultra-long roaming, allowing for rapid generation of interactive and immersive 3D scenes based on a single image or text.

TranslateGemma
Google's open source lightweight multimodal translation model supports 55 languages and image translations, with performance that exceeds larger models, taking into account both mobile and cloud deployments, and facilitating efficient globalized communication.

NVIDIA Ising
The world's first open-source quantum AI model series, through AI-driven quantum chip calibration and error correction, provides a high-performance tool chain for practical quantum computing and reshapes the quantum industry ecosystem.

Gemini 3
Google launched the world's first native multimodal “doctoral” AI model, with millions of contexts, cross-modal deep reasoning and generative UI as the core, redefining the boundaries of intelligent collaboration from scientific research and creation to everyday tasks.

BLOOM
A large open-source multilingual language model developed by over 1,000 researchers from more than 60 countries and 250 institutions, with 176B parameters and trained on the ROOTS corpus, supporting 46 natural languages and 13 programming languages, aims to advance the research and use of large-scale language models by academics and small companies.

Phi-3
A high-performance large-scale language model from Microsoft, tuned with instructions to support cross-platform operation, with excellent language comprehension and reasoning capabilities, especially suitable for multimodal application scenarios.

Qwen3-Coder
Ali open source code big model, support full-flow programming and complex task planning, performance over GPT-4.1, lower cost.

OmAgent
Device-oriented open-source smart body framework designed to simplify the development of multimodal smart bodies and provide enhancements for various types of hardware devices.
No comments...
