Categories: FAANG

Matrix3D: Large Photogrammetry Model All-in-One

We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs…
AI Generated Robotic Content

Recent Posts

TaoMate – H3 3 steps lora used as a refiner

The lora itself at 3 steps is nothing to write home about. If the scene…

1 hour ago

AI uncovers hidden Ozempic side effects across 400,000 Reddit posts

AI analysis of 400,000 Reddit posts found that users of drugs such as Ozempic, Wegovy,…

2 hours ago

Anthropic CEO Dario Amodei says AI industry needs to give safety measures time to catch up

The CEO of Anthropic said Saturday the artificial-intelligence industry should slow its fast-moving development to…

2 hours ago

FLUX.2-klein-9B RefMods

submitted by /u/malcolmrey [link] [comments]

1 day ago

10 Best Standing Desks Worth Buying in 2026

Take your home office to new heights with our favorite motorized standing desks.

1 day ago

Testing MiniMax-H3 Physics knowledge Pt2

Some weeks ago, I posted a set of experiments to "understand" the physical knowledge of…

2 days ago