
What is LingBot-Depth 2.0?
LingBot-Depth 2.0 is a next-generation foundational model for robotic spatial perception released on July 7, 2026, by Robbyant, a subsidiary of Ant Group. It is not positioned as a large language model, but rather as a foundation model for 3D vision, primarily designed to enable robots to ”see the three-dimensional world clearly.”
GPT is the brain of AI, while LingBot-Depth is the robot's eyes.It addresses one of the most fundamental issues in robotics—Depth Perception.
Technical Principles of LingBot-Depth 2.0
The breakthroughs achieved by LingBot-Depth 2.0 are primarily attributable to the following two core technologies:
- Spatial Native Vision Platform (LingBot-Vision): It is the industry’s first visual foundation model to use “boundary structure” as a pre-training objective. By introducing geometric constraints and a boundary-centered masking mechanism, it possesses sub-pixel-level boundary localization and spatial structure understanding capabilities, providing robust foundational support for depth estimation.
- Masked Deep Modeling (MDM) and Big Data: The model uses the “depth gaps” naturally created by depth sensors on transparent and reflective objects as masks, requiring the model to “fill in” the complete depth using only RGB color images and the remaining valid depth information. At the same time, the scale of its training data has expanded significantly from 3 million in version 1.0 to 150 million, creating a virtuous cycle between data and the model.
Technological Advantage:
| abilities | LingBot-Depth 2.0 |
|---|---|
| RGB + Depth Fusion | ✅ |
| Metric Depth | ✅ |
| Advanced Autocomplete | ⭐⭐⭐⭐⭐ |
| Transparent Objects | ⭐⭐⭐⭐⭐ |
| Mirror | ⭐⭐⭐⭐⭐ |
| Small Object Recognition | ⭐⭐⭐⭐☆ |
| Edge Sharpness | ⭐⭐⭐⭐⭐ |
| Robotic Gripping | ⭐⭐⭐⭐⭐ |
| 3D reconstruction | ⭐⭐⭐⭐⭐ |
| SLAM | ⭐⭐⭐⭐☆ |
Key Features of LingBot-Depth 2.0
- In-Depth Autocomplete for Complex Materials: It effectively overcomes the “hard scenes”—such as glass, mirrors, and transparent objects—where traditional depth cameras often fail, and is capable of generating 3D depth maps with complete structures and clear boundaries.
- Precise Spatial Awareness: It delivers outstanding performance in edge clarity, recognition of small objects (such as cables and thin rods), and long-range depth estimation, thereby preventing obstacle-avoidance blind spots caused by the robot missing objects.
- High Precision and High Robustness: In the most challenging indoor scenarios involving large-scale depth loss, the depth error (RMSE) was halved compared to the previous generation (from 0.132 to 0.062), and the system ranked first in 12 out of 16 depth completion benchmark evaluations.
Use Cases for LingBot-Depth 2.0
- Embodied Intelligence and Robotics: Provides a reliable spatial data foundation for robot navigation, obstacle avoidance, and precise grasping (such as grasping transparent storage bins, glass cups, and champagne towers).
- industrial automation: Enables tasks such as precision assembly, alignment of highly reflective metal parts, and tracking the edges of cardboard boxes in logistics and warehousing.
- AR/VR and Autonomous Driving: Provides precise depth information to enhance safety and user experience in environments with complex lighting and materials.
LingBot-Depth 2.0 project repository
- Project website::https://technology.robbyant.com/lingbot-vision
- GitHub repository::https://github.com/robbyant/lingbot-vision
- HuggingFace Model Library::https://huggingface.co/collections/robbyant/lingbot-vision
- arXiv Technical Paper::https://github.com/robbyant/lingbot-vision/blob/main/paper.pdf
How do I use LingBot-Depth 2.0?
- Get the source code: The model weights for LingBot-Vision (available in four versions: ViT-G, L, B, and S) and its technical report have been fully open-sourced; developers can access them via GitHub, Hugging Face, or ModelScope.
- Hardware and SDK Integration: Ant Lingbo has entered into a deep partnership with Orbbec. Users can enable spatial perception capabilities with a single click using Orbbec’s SDK, or plan to purchase an all-in-one “3D camera + spatial perception” camera product—which integrates the commercial version of this model—by the end of the year for out-of-the-box use.
- Cloud Deployment: Developers can also leverage AI computing platforms such as Baidu Baige to quickly run models and perform inference using the official Docker images provided.
Comparison of similar products
| offerings | firms | localization | dominance | inferior |
|---|---|---|---|---|
| LingBot-Depth 2.0 | Ant Lingbo | Spatial Perception Foundation Model | Outstanding performance with transparent objects, mirrors, and depth restoration | The ecosystem is still under development |
| Intel RealSense SDK | Intel | Deep Learning Algorithms | Mature and stable | The glass performed averagely |
| Orbbec SDK | Aobi Zhongguang | Camera Algorithms | Domestic Ecosystem | More reliance on hardware |
| NVIDIA Isaac Perceptor | NVIDIA | Robotic Vision | Seamless integration with the Isaac platform | Heavily reliant on the NVIDIA ecosystem |
| Meta SAM2 | Meta | Image Segmentation | Strong segmentation capabilities | Does not provide genuine depth |
| Depth Anything V2 | Depth Estimation | Monocular Depth | No depth camera required | Not Actual Scale (Metric Depth) |
| Metric3D | Tsinghua University, etc. | Monocular Metric Depth | High accuracy | Limited support for transparent objects |
| FoundationStereo | Meta, etc. | Stereoscopic Depth | Both eyes are estimated to be excellent | Complex materials still pose a challenge |
The biggest difference is:
LingBot-Depth is not a monocular depth estimation model, but rather a ”depth completion + depth enhancement” model designed for RGB-D cameras, with a greater emphasis on spatial perception in scenes featuring real-world scales and complex materials.
data statistics
Relevant Navigation

Wanxing Technology has developed China's first audio and video multimedia creation pendant big model, which integrates video, audio, picture and language processing capabilities to provide powerful AI creation support for the digital creative field.

SmartResume
Ali open source SmartResume is a high-precision resume parsing system based on OCR and lightweight large models, which can convert 12 formats of resumes such as PDF/pictures into structured data in seconds, with an accuracy rate of 93.1%.

DeepSeek
Developed by Hangzhou Depth Seeker, a large open source AI project integrating natural language processing and code generation capabilities, supporting efficient information search and answering services.

Gemini 2.0 Flash
Google introduced a new generation of AI models that support multimodal inputs and outputs and natively integrate intelligent tools to provide developers with powerful and flexible assistant functions.

Manifold AI
Focusing on the independent development of an embodied intelligence world model, the company creates a "digital brain" for robots that understands the laws of physics and supports real-time interaction. It is the first company in China to directly apply a world model to robotic applications.

BLOOM
A large open-source multilingual language model developed by over 1,000 researchers from more than 60 countries and 250 institutions, with 176B parameters and trained on the ROOTS corpus, supporting 46 natural languages and 13 programming languages, aims to advance the research and use of large-scale language models by academics and small companies.

ZhiPu AI BM
The series of large models jointly developed by Tsinghua University and Smart Spectrum AI have powerful multimodal understanding and generation capabilities, and are widely used in natural language processing, code generation and other scenarios.

OmniGen
Unified image generation diffusion model, which naturally supports multiple image generation tasks with high flexibility and scalability.
No comments...
