Junior Assistant Professor · University of Bologna
Fabio Tosi, PhD
I work at the intersection of computer vision and deep learning, on machines that understand the 3D structure of the world — stereo matching, monocular depth estimation, neural rendering and SLAM.
Department of Computer Science and Engineering (DISI), University of Bologna. I am also interested in making these models small and fast enough to run on resource-constrained devices.
About
- Stereo Matching
- Monocular Depth Estimation
- Neural Rendering
- SLAM
- Efficient Deep Learning
I received my PhD in 2021 from the University of Bologna, supervised by Professor Stefano Mattoccia, and my Master’s (2017) and Bachelor’s (2014) degrees in Computer Engineering from the same university.
In 2020, I was a visiting PhD student in the Autonomous Vision Group (AVG) led by Professor Andreas Geiger at the Max Planck Institute for Intelligent Systems and the University of Tübingen. In 2022, I received the Best PhD Thesis Award from the Italian Association for Research in Computer Vision, Pattern Recognition and Machine Learning (CVPL).
Since 2024, I serve as Associate Editor for Pattern Recognition (Elsevier). I have also served as Area Chair at CVPR 2026, ECCV 2026, ACCV 2026, and ICIAP 2025, and as Associate Editor at IROS 2025.
News
- 09/2026 Honored to serve again as Area Chair at CVPR 2027!
- 09/2026 Our paper “Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation” has been accepted at SIGGRAPH Asia 2026 (ACM Transactions on Graphics)!
- 09/2026 The 3rd Workshop on Neural SLAM (NeuSLAM) I co-organized took place at ECCV 2026 in Malmö, Sweden. Thanks to all the speakers and everyone who joined!
- 08/2026 1 paper accepted to BMVC 2026!
- 06/2026 3 papers accepted to ECCV 2026, 1 paper accepted to IROS 2026!
- 05/2026 Honored to be recognized as an Outstanding Area Chair at CVPR 2026!
- 02/2026 3 papers at CVPR 2026, 1 at CVPR Findings 2026!
- 02/2026 Our survey on NeRF & 3DGS-based SLAM accepted to T-RO!
Earlier newsHide earlier news
- 12/2025 2 paper accepted to WACV 2025 and AAAI 2025 (Oral)!
- 09/2024 Grateful for the Outstanding Reviewer recognition at ICCV 2025!
- 09/2025 1 paper accepted with Rawmantic AI to NIPS 2025!
- 06/2025 2 paper accepted to ICCV 2025!
- 06/2025 One paper is set for IROS 2025, and another for IJCV!
- 05/2025 Grateful for the recognition as an Outstanding Reviewer at CVPR 2025!
- 02/2025 Our paper “Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail” has been accepted at CVPR 2025!
- 01/2025 1 paper accepted to ICRA 2025!
- 12/2024 Our survey paper, “A Survey on Deep Stereo Matching in the Twenties”, has been accepted at the IJCV journal!
- 11/2024 Our research with Sadra and Fatma received the Best Poster Award at BMVC 2024!
- 07/2024 1 paper accepted to WACV 2024!
- 09/2024 Grateful for the Outstanding Reviewer recognition at ECCV 2024!
- 07/2024 Our paper “Diffusion Models for Monocular Depth Estimation” has been accepted at ECCV 2024!
- 06/2024 Our CVPR tutorial on deep stereo matching is now available! Get all the insights here!
- 06/2024 Our extended version of Neural Disparity Refinement has been accepted for publication in the TPAMI journal!
- 05/2024 Honored to be recognized as a Outstanding Reviewer at CVPR 2024!
- 10/2023 I am proud to announce my new role as a Junior Assistant Professor (RTDA) at the Department of Computer Science and Engineering (DISI)!
- 05/2023 It is with great pleasure that I announce my achievement as an Outstanding Reviewer at CVPR 2023! (Award Certificate)
- 02/2023 I received my National Scientific Habilitation (09/H1)
- 11/2022 Celebrating the recognition of the Best PhD Thesis Award from the Italian Association for Computer Vision Research (CVPL 2022)
- 09/2021 Best Paper Honorable Mention to our work “Neural Disparity Refinement for Arbitrary Resolution Stereo”
Teaching
University of Bologna
Accelerated Computing Systems
Sistemi di Elaborazione Accelerata M
CUDA and the GPU software stack, GPU architectures and high-performance computing: how to write code that actually keeps a modern GPU busy. Module 2, alongside Stefano Mattoccia.
Fundamentals of Computer Science
Logic networks and computer architectures — how a machine gets from gates to instructions.
GPU-accelerated Computing for AI
The same machinery seen from the side of deep learning: where the time really goes when a model trains, and what can be done about it.
What we work on
Recent results from the group — click to filter the publications below
Thesis & internships
For MSc students in Computer Engineering and Artificial Intelligence at the University of Bologna
If you are looking for a thesis that is an open research problem rather than a closed exercise, get in touch. You would work inside our group at CVLab, with our GPUs, our codebases and regular supervision, on a topic close to what we publish — and the strongest projects can grow into a paper.
Monocular depth estimation
Depth foundation models: sharper boundaries, robustness to hard conditions, and making them small and fast enough to run on a phone or an embedded board.
Stereo matching
Zero-shot generalization, transparent and reflective surfaces, event cameras, and stereo networks that adapt on the fly to the scene in front of them.
Multi-view stereo & 3D reconstruction
Feed-forward reconstruction from casual video, neural rendering and 3D Gaussian Splatting, dense SLAM built on depth foundation models.
Vision-Language Models New
Grounding 3D perception in language: what a VLM can and cannot say about geometry, and how spatial understanding can be taught to one.
Vision-Language-Action New
From perception to action: how far accurate 3D perception takes a VLA policy, and where robot manipulation still breaks.
Your own idea
If you have a proposal that overlaps with what we do, bring it. The topics above are where we are strongest, not a closed list.
A period abroad. We collaborate with universities and companies worldwide, and students from the group have already spent research periods abroad. For students who are doing well, we can explore a visit or an internship with one of our partners — it is not guaranteed, but it is worth asking about early.
Write to me with your CV, your transcript, and a few lines about what you find interesting and what you would like to learn.
Get in touchSelected publications
A selection — the full list is on Google Scholar · * indicates joint first authorship

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation New
SIGGRAPH Asia|ACM Transactions on Graphics|vol. 45, no. 6, art. 204
project pagearXivcodeweightsdemo
TL;DR
Turns a pretrained image-editing diffusion transformer into a single-step depth estimator, through representation alignment and a two-stage fine-tuning built on a Sinkhorn loss. It recovers fur, foliage and hair-thin edges that earlier models smooth away, while staying cheap enough to run.

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device New
ECCV|A++|European Conference on Computer Vision
project pagearXivsupplementarycodedemovideo
TL;DR
A 6.1M-parameter depth network distilled from a foundation model over many domains. It keeps zero-shot generalization while running in real time on embedded hardware, closing much of the gap to models fifty times larger.

DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations New
ECCV|A++|European Conference on Computer Vision
TL;DR
Feeds DINO features, enriched by a scene geometry encoder, into NeRF- and Gaussian-Splatting SLAM. The map then carries semantics and geometry together, instead of appearance alone.

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos New
ECCV|A++|European Conference on Computer Vision
TL;DR
Reconstructs and tracks from monocular video at nearly 10 FPS by splitting the work across agents: each regresses local point maps, a merging model fuses them, and pose-graph optimization cancels the drift a feed-forward pipeline accumulates.

FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow New
BMVC|A|British Machine Vision Conference
TL;DR
Treats optical flow as global matching. A hierarchical transformer supplies long-range context, and casting the initialization as an optimal transport problem yields a robust starting flow together with explicit occlusion and confidence maps, which then guide the refinement.

Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
CVPR|A++|Conference on Computer Vision and Pattern Recognition
TL;DR
Stereo between an event camera and an ordinary one. Prompting the two modalities in both directions keeps the cues specific to each from being washed out by the other, which is what usually breaks asymmetric stereo.

EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors
CVPR|A++|Conference on Computer Vision and Pattern Recognition
project pagepaperarXivpostercode
TL;DR
Trains event-based stereo without active depth sensors. Novel view synthesis turns ordinary colour images into proxy events and proxy labels, and stereo models from the RGB literature are repurposed on that data, generalizing far beyond what annotated event datasets allow.

StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
CVPR Findings|Conference on Computer Vision and Pattern Recognition – Findings
TL;DR
Synthesizes the second view of a stereo pair directly with diffusion, conditioned on viewpoint in a canonical rectified space — no depth estimation, no warping. It also proposes an evaluation protocol that forbids ground-truth geometry at test time.

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
CVPR|A++|Conference on Computer Vision and Pattern Recognition
TL;DR
Open-vocabulary semantic 3D reconstruction from RGB video: CLIP semantics enter the reconstruction itself rather than being painted on afterwards, and 2D features are lifted into 3D descriptors that fuse space, geometry and meaning.

How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey
T-RO|Q1 · IF 10.8|IEEE Transactions on Robotics|vol. 42, pp. 1405–1427
TL;DR
The first survey of SLAM seen through radiance fields: how NeRF and 3D Gaussian Splatting reshaped mapping and tracking, what each representation buys you, and where they still break.

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM Oral
AAAI|A++|AAAI Conference on Artificial Intelligence
TL;DR
Monocular dense SLAM that grounds flow estimation in depth foundation models, so correspondences stay geometrically consistent, and then enforces global consistency with a bundle adjustment layer optimizing poses and depth jointly.

WarpRF: Multi-View Consistency for Training-Free Uncertainty Quantification and Applications in Radiance Fields
WACV|A|IEEE/CVF Winter Conference on Applications of Computer Vision
TL;DR
Measures how much a radiance field can be trusted without training anything. Render from the views you have, warp them into one you do not, and see whether the model agrees with itself.

Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting
NeurIPS|A++|Conference on Neural Information Processing Systems
TL;DR
Optimizes 3D Gaussian Splatting and the vision-model priors that supervise it at the same time, so each keeps improving the other, with a bundle-adjustment step that escapes the purely local supervision of standard 3DGS pipelines.

FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
ICCV|A++|International Conference on Computer Vision
TL;DR
Optical flow trained on a single consumer GPU — roughly eight times less hardware than comparable methods — by pairing depth foundation models with a classical low-dimensional motion parametrization, and still generalizing better across datasets.

A Survey on Deep Stereo Matching in the Twenties
IJCV|Q1 · IF 11.6|International Journal of Computer Vision
TL;DR
Maps the 2020s of deep stereo: the architectures and paradigms that redefined the field in the last five years, the challenges that stayed open, and a quantitative account of where the benchmarks now stand.

Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail
CVPR|A++|Conference on Computer Vision and Pattern Recognition
TL;DR
Couples a stereo network with monocular priors from a vision foundation model, so the two cover each other's blind spots. Trained only on synthetic data, it still holds up on mirrors, transparencies and textureless regions, where either cue alone fails.

Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation
ICCV|A++|International Conference on Computer Vision
TL;DR
Brings monocular depth to event cameras, where dense ground truth essentially does not exist, by distilling a vision foundation model through spatially aligned RGB into dense proxy labels for the event stream.

Active Stereo in the Wild through Virtual Pattern Projection
IJCV|Q1 · IF 11.6|International Journal of Computer Vision
TL;DR
Replaces the physical projector of active stereo with a virtual one: sparse measurements from any depth sensor are painted onto both images as patterns consistent with the scene, turning a passive rig into an active one without the projector's range and lighting limits.

CabNIR: A Benchmark for In-Vehicle Infrared Monocular Depth Estimation
WACV|A|Winter Conference on Applications of Computer Vision
TL;DR
A near-infrared benchmark for depth estimation inside the cabin: more than 41,000 frames with ground truth, across 36 vehicles and 45 participants — the scale in-vehicle depth research was missing.

HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM
ICRA|A|International Conference on Robotics and Automation
TL;DR
NeRF-based SLAM with a hybrid encoding — hash grid, tri-planes and one-blob together — plus structural supervision, aimed at the scenes where existing systems lose completeness or drift out of global consistency.

Self-Evolving Depth-Supervised 3D Gaussian Splatting from Rendered Stereo Pairs 🏆 Best Poster Award
BMVC|A|British Machine Vision Conference
TL;DR
The Gaussian Splatting model renders virtual stereo pairs of itself, a stereo network turns them into depth supervision, and that supervision repairs the floating artifacts in its own geometry. A loop that improves as it runs.

Diffusion Models for Monocular Depth Estimation: Overcoming Challenging Conditions
ECCV|A++|European Conference on Computer Vision
TL;DR
Generates the hard cases instead of hunting for them: text-to-image diffusion with depth-aware control turns easy scenes into rainy, dark or otherwise adverse ones while preserving their depth, so monocular networks can learn conditions nobody has labelled.

Booster: a Benchmark for Depth from Images of Specular and Transparent Surfaces
TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence|vol. 46, no. 1, pp. 85–102
TL;DR
Dense, high-resolution ground truth for exactly what breaks depth estimation — mirrors, glass and other non-Lambertian surfaces — labelled with sub-pixel precision through a deep space-time stereo pipeline.

Neural Disparity Refinement
TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence
TL;DR
The journal version of neural disparity refinement: a continuous formulation that outputs a refined disparity map at any resolution, built for phones, where a high-resolution and a low-resolution camera have to cooperate.

Federated Online Adaptation for Deep Stereo
CVPR|A++|Conference on Computer Vision and Pattern Recognition
TL;DR
Stereo networks deployed in different environments share what they learn while adapting online. A device that cannot afford to adapt on its own still benefits from the experience of the others.

GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction
ICCV|A++|International Conference on Computer Vision
TL;DR
Dense neural SLAM that keeps poses and reconstruction globally consistent in real time, through loop closing and online full bundle adjustment, instead of letting tracking error accumulate into a distorted map.

Active Stereo Without Pattern Projector
ICCV|A++|International Conference on Computer Vision
TL;DR
Gives a passive stereo pair the benefits of active stereo with no projector at all: sparse depth hints are virtually projected onto both images as a pattern, so correspondence becomes easy where texture is missing.

NeRF-Supervised Deep Stereo
CVPR|A++|Conference on Computer Vision and Pattern Recognition
project pagepapersupplementarycodedatasetvideo
TL;DR
Trains stereo networks with no ground truth whatsoever. A handheld video becomes a NeRF, the NeRF renders stereo triplets and proxy depth, and the network learns sharp disparities from images that were never captured by a stereo rig.

Learning Depth Estimation for Transparent and Mirror Surfaces
ICCV|A++|International Conference on Computer Vision
TL;DR
Teaches depth for transparent and mirror surfaces without a single annotation: in-paint the offending object, let a monocular model label the repaired image, and fine-tune on those pseudo labels.

GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor Scenes
ICCV|A++|International Conference on Computer Vision
TL;DR
Self-supervised indoor depth, where large rotations and bare walls break the usual recipe. Coarse poses from multi-view geometry help, but only once the scale ambiguity across scenes is handled — which is the paper's actual contribution.

MonoViT: Self-supervised Monocular Depth Estimation with a Vision Transformer
3DV|A-|International Conference on 3D Vision
TL;DR
Brings the global reasoning of vision transformers to self-supervised monocular depth, where the limited receptive field of convolutions had confined the network to local decisions.

Cross-Spectral Neural Radiance Fields
3DV|A-|International Conference on 3D Vision
papervideo & supplementarydataset
TL;DR
A radiance field shared by cameras that see different parts of the spectrum — colour, multispectral, infrared. It optimizes poses across spectra so that any viewpoint can be rendered in any modality, aligned and at the same resolution.

Open Challenges in Deep Stereo: the Booster Dataset
CVPR|A++|Conference on Computer Vision and Pattern Recognition
papersupplementarydatasetbenchmarkvideo
TL;DR
419 high-resolution indoor samples across 64 scenes, densely annotated and deliberately full of specular and transparent surfaces: a benchmark built around the cases where stereo networks fail.

RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation
CVPR|A++|Conference on Computer Vision and Pattern Recognition
papersupplementarydatasetvideo
TL;DR
Registers colour and multispectral images of very different resolution by treating it as stereo matching, with a new dataset and an architecture trained self-supervised by borrowing a third camera as supervision.

Continual Adaptation for Deep Stereo
TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence|vol. 44, no. 9, pp. 4713–4729
TL;DR
Adaptation that never stops: rather than assuming the training distribution covers deployment, the stereo network keeps adjusting to whatever environment it actually meets.

On the Confidence of Stereo Matching in a Deep-Learning Era: A Quantitative Evaluation
TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence|vol. 44, no. 9, pp. 5293–5313
TL;DR
A quantitative account of how far confidence estimation for stereo has come in the deep learning era — which measures are worth trusting, and what that reliability buys the algorithms downstream.

Neural Disparity Refinement for Arbitrary Resolution Stereo 🏆 Best Paper Honorable Mention
3DV|A-|International Conference on 3D Vision
project pagepaper & supplementarycode
TL;DR
Refines a disparity map at any output resolution through a continuous formulation, so cheap consumer devices — phones with one high-resolution and one low-resolution camera — can produce clean 3D.

On the Synergies Between Machine Learning and Binocular Stereo for Depth Estimation From Images: A Survey
TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence
TL;DR
Forty years of stereo read through the lens of machine learning: how the two traditions met, where learning genuinely helped, and which problems survived the transition.

SMD-Nets: Stereo Mixture Density Networks
CVPR|A++|Conference on Computer Vision and Pattern Recognition
papersupplementblogcodevideoposter
TL;DR
Predicts a bimodal mixture density instead of a single disparity per pixel. Edges stay sharp where depth jumps, and the output can be sampled at arbitrary resolution.

Distilled Semantics for Comprehensive Scene Understanding from Videos
CVPR|A++|Conference on Computer Vision and Pattern Recognition
TL;DR
Learns depth, motion and semantics together from monocular video, with the semantic supervision distilled from a pretrained network — no manual labels for any of the three.

Reversing the Cycle: Self-Supervised Deep Stereo through Enhanced Monocular Distillation
ECCV|A++|European Conference on Computer Vision
TL;DR
Reverses the usual direction of self-supervision: instead of stereo teaching a monocular network, a monocular completion network is distilled into a stereo one, which softens the artifacts stereo self-supervision leaves behind.

Learning Monocular Depth Estimation Infusing Traditional Stereo Knowledge
CVPR|A++|Conference on Computer Vision and Pattern Recognition
papersupplementarycodepostervideo
TL;DR
Infers depth from a single image by synthesizing the features of a second view and matching them, importing the machinery of stereo into a monocular network — and training it without any labels, using traditional stereo as the teacher.

Real-Time Self-Adaptive Deep Stereo Oral
CVPR|A++|Conference on Computer Vision and Pattern Recognition
papersupplementarycodevideolive demo
TL;DR
A stereo network that fine-tunes itself online while it runs, so accuracy does not collapse the moment the scene stops resembling the training set.

Guided Stereo Matching
CVPR|A++|Conference on Computer Vision and Pattern Recognition
TL;DR
A handful of sparse but reliable depth measurements, fed to the network at inference, steer stereo matching back to accuracy when the environment changes and a fixed model would drift.

Quantitative Evaluation of Confidence Measures in a Machine Learning World Spotlight
ICCV|A++|International Conference on Computer Vision
TL;DR
An extensive comparison of confidence measures for stereo, hand-crafted and learned, across algorithms and datasets — the reference point for telling a good match from a bad one.







