Fabio Tosi

Junior Assistant Professor · University of Bologna

Fabio Tosi, PhD

I work at the intersection of computer vision and deep learning, on machines that understand the 3D structure of the world — stereo matching, monocular depth estimation, neural rendering and SLAM.

Department of Computer Science and Engineering (DISI), University of Bologna. I am also interested in making these models small and fast enough to run on resource-constrained devices.

Fabio Tosi
Drag to compare:

Depth and normals estimated with Marigold V2
Bologna, Italy · fabio.tosi5@unibo.it

About

Research
  • Stereo Matching
  • Monocular Depth Estimation
  • Neural Rendering
  • SLAM
  • Efficient Deep Learning
Background

I received my PhD in 2021 from the University of Bologna, supervised by Professor Stefano Mattoccia, and my Master’s (2017) and Bachelor’s (2014) degrees in Computer Engineering from the same university.

In 2020, I was a visiting PhD student in the Autonomous Vision Group (AVG) led by Professor Andreas Geiger at the Max Planck Institute for Intelligent Systems and the University of Tübingen. In 2022, I received the Best PhD Thesis Award from the Italian Association for Research in Computer Vision, Pattern Recognition and Machine Learning (CVPL).

Service

Since 2024, I serve as Associate Editor for Pattern Recognition (Elsevier). I have also served as Area Chair at CVPR 2026, ECCV 2026, ACCV 2026, and ICIAP 2025, and as Associate Editor at IROS 2025.

News

  • 09/2026 Honored to serve again as Area Chair at CVPR 2027!
  • 09/2026 Our paper “Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation” has been accepted at SIGGRAPH Asia 2026 (ACM Transactions on Graphics)!
  • 09/2026 The 3rd Workshop on Neural SLAM (NeuSLAM) I co-organized took place at ECCV 2026 in Malmö, Sweden. Thanks to all the speakers and everyone who joined!
  • 08/2026 1 paper accepted to BMVC 2026!
  • 06/2026 3 papers accepted to ECCV 2026, 1 paper accepted to IROS 2026!
  • 05/2026 Honored to be recognized as an Outstanding Area Chair at CVPR 2026!
  • 02/2026 3 papers at CVPR 2026, 1 at CVPR Findings 2026!
  • 02/2026 Our survey on NeRF & 3DGS-based SLAM accepted to T-RO!
Earlier newsHide earlier news

Research team

CVLab — University of Bologna

Stefano Mattoccia Stefano Mattoccia Full Professor
Matteo Poggi Matteo Poggi Associate Professor
Luca Bartolomei Luca Bartolomei Post-doc
Ziren Gong Ziren Gong PhD Student
Ugo Leone Cavalcanti Ugo Leone Cavalcanti PhD Student
Enrico Mannocci Enrico Mannocci PhD Student

Teaching

University of Bologna

  • Accelerated Computing Systems

    Sistemi di Elaborazione Accelerata M

    CUDA and the GPU software stack, GPU architectures and high-performance computing: how to write code that actually keeps a modern GPU busy. Module 2, alongside Stefano Mattoccia.

    MSc Computer Engineering (cod. 6719) B8563 · 8 CFU · taught in Italian
  • Fundamentals of Computer Science

    Logic networks and computer architectures — how a machine gets from gates to instructions.

    BSc Mechatronics (cod. 6009) B0041 · 3 CFU · taught in Italian
  • GPU-accelerated Computing for AI

    The same machinery seen from the side of deep learning: where the time really goes when a model trains, and what can be done about it.

    PhD Doctoral courses taught in English

What we work on

Recent results from the group — click to filter the publications below

Thesis & internships

For MSc students in Computer Engineering and Artificial Intelligence at the University of Bologna

If you are looking for a thesis that is an open research problem rather than a closed exercise, get in touch. You would work inside our group at CVLab, with our GPUs, our codebases and regular supervision, on a topic close to what we publish — and the strongest projects can grow into a paper.

Monocular depth estimation

Depth foundation models: sharper boundaries, robustness to hard conditions, and making them small and fast enough to run on a phone or an embedded board.

Stereo matching

Zero-shot generalization, transparent and reflective surfaces, event cameras, and stereo networks that adapt on the fly to the scene in front of them.

Multi-view stereo & 3D reconstruction

Feed-forward reconstruction from casual video, neural rendering and 3D Gaussian Splatting, dense SLAM built on depth foundation models.

Vision-Language Models New

Grounding 3D perception in language: what a VLM can and cannot say about geometry, and how spatial understanding can be taught to one.

Vision-Language-Action New

From perception to action: how far accurate 3D perception takes a VLA policy, and where robot manipulation still breaks.

Your own idea

If you have a proposal that overlaps with what we do, bring it. The topics above are where we are strongest, not a closed list.

A period abroad. We collaborate with universities and companies worldwide, and students from the group have already spent research periods abroad. For students who are doing well, we can explore a visit or an internship with one of our partners — it is not guaranteed, but it is worth asking about early.

Write to me with your CV, your transcript, and a few lines about what you find interesting and what you would like to learn.

Get in touch

Selected publications

A selection — the full list is on Google Scholar · * indicates joint first authorship

2026
Marigold V2

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation New

Igor Pavlovic*, Thiemo Wandel*, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai

SIGGRAPH Asia|ACM Transactions on Graphics|vol. 45, no. 6, art. 204

TL;DR

Turns a pretrained image-editing diffusion transformer into a single-step depth estimator, through representation alignment and a two-stage fine-tuning built on a Sinkhorn loss. It recovers fur, foliage and hair-thin edges that earlier models smooth away, while staying cheap enough to run.

2026
ZipDepth

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device New

Fabio Tosi, Luca Bartolomei, Matteo Poggi, Stefano Mattoccia

ECCV|A++|European Conference on Computer Vision

TL;DR

A 6.1M-parameter depth network distilled from a foundation model over many domains. It keeps zero-shot generalization while running in real time on embedded hardware, closing much of the gap to models fifty times larger.

2026
DINO-SLAM

DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations New

Ziren Gong, Xiaohan Li, Fabio Tosi, Youmin Zhang, Stefano Mattoccia, Jiawei Wu, Matteo Poggi

ECCV|A++|European Conference on Computer Vision

TL;DR

Feeds DINO features, enriched by a scene geometry encoder, into NeRF- and Gaussian-Splatting SLAM. The map then carries semantics and geometry together, instead of appearance alone.

2026
MAGiSt3R

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos New

Ziren Gong, Xiaohan Li, Fabio Tosi, Ninghui Xu, Stefano Mattoccia, Jianfei Cai, Matteo Poggi

ECCV|A++|European Conference on Computer Vision

TL;DR

Reconstructs and tracks from monocular video at nearly 10 FPS by splitting the work across agents: each regresses local point maps, a merging model fuses them, and pose-graph optimization cancels the drift a feed-forward pipeline accumulates.

2026
FlowIt

FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow New

Sadra Safadoust, Fabio Tosi, Matteo Poggi, Fatma Güney

BMVC|A|British Machine Vision Conference

TL;DR

Treats optical flow as global matching. A hierarchical transformer supplies long-range context, and casting the initialization as an optimal transport problem yields a robust starting flow together with explicit occlusion and confidence maps, which then guide the refinement.

2026
Bi-CMPStereo

Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

Ninghui Xu, Fabio Tosi, Lihui Wang, Jiawei Han, Luca Bartolomei, Zhiting Yao, Matteo Poggi, Stefano Mattoccia

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Stereo between an event camera and an ordinary one. Prompting the two modalities in both directions keeps the cues specific to each from being washed out by the other, which is what usually breaks asymmetric stereo.

2026
EventHub

EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors

Luca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano Mattoccia, Guillermo Gallego

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Trains event-based stereo without active depth sensors. Novel view synthesis turns ordinary colour images into proxy events and proxy labels, and stereo models from the RGB literature are repurposed on that data, generalizing far beyond what annotated event datasets allow.

2026
StereoSpace

StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

Tjark Behrens, Anton Obukhov, Bingxin Ke, Fabio Tosi, Matteo Poggi, Konrad Schindler

CVPR Findings|Conference on Computer Vision and Pattern Recognition – Findings

TL;DR

Synthesizes the second view of a stereo pair directly with diffusion, conditioned on viewpoint in a canonical rectified space — no depth estimation, no warping. It also proposes an evaluation protocol that forbids ground-truth geometry at test time.

2026
Ov3R

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

Ziren Gong, Xiaohan Li, Fabio Tosi, Jiawei Han, Stefano Mattoccia, Jianfei Cai, Matteo Poggi

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Open-vocabulary semantic 3D reconstruction from RGB video: CLIP semantics enter the reconstruction itself rather than being painted on afterwards, and 2D features are lifted into 3D descriptors that fuse space, geometry and meaning.

2026
NeRF and 3DGS SLAM survey

How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey

Fabio Tosi, Youmin Zhang, Ziren Gong, Erik Sandström, Stefano Mattoccia, Martin R. Oswald, Matteo Poggi

T-RO|Q1 · IF 10.8|IEEE Transactions on Robotics|vol. 42, pp. 1405–1427

TL;DR

The first survey of SLAM seen through radiance fields: how NeRF and 3D Gaussian Splatting reshaped mapping and tracking, what each representation buys you, and where they still break.

2026
FoundationSLAM

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM Oral

Yuchen Wu, Jiahe Li, Fabio Tosi, Matteo Poggi, Jin Zheng, Xiao Bai

AAAI|A++|AAAI Conference on Artificial Intelligence

TL;DR

Monocular dense SLAM that grounds flow estimation in depth foundation models, so correspondences stay geometrically consistent, and then enforces global consistency with a bundle adjustment layer optimizing poses and depth jointly.

2026
WarpRF

WarpRF: Multi-View Consistency for Training-Free Uncertainty Quantification and Applications in Radiance Fields

Sadra Safadoust, Fabio Tosi, Fatma Güney, Matteo Poggi

WACV|A|IEEE/CVF Winter Conference on Applications of Computer Vision

TL;DR

Measures how much a radiance field can be trusted without training anything. Render from the views you have, warp them into one you do not, and see whether the model agrees with itself.

2025
Eve3D

Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting

Jiawei Zhang, Youmin Zhang, Fabio Tosi, Meiying Gu, Jiahe Li, Xiaohan Yu, Jin Zheng, Xiao Bai, Matteo Poggi

NeurIPS|A++|Conference on Neural Information Processing Systems

TL;DR

Optimizes 3D Gaussian Splatting and the vision-model priors that supervise it at the same time, so each keeps improving the other, with a bundle-adjustment step that escapes the purely local supervision of standard 3DGS pipelines.

2025
FlowSeek

FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases

Matteo Poggi, Fabio Tosi

ICCV|A++|International Conference on Computer Vision

TL;DR

Optical flow trained on a single consumer GPU — roughly eight times less hardware than comparable methods — by pairing depth foundation models with a classical low-dimensional motion parametrization, and still generalizing better across datasets.

2025
A Survey on Deep Stereo Matching in the Twenties

A Survey on Deep Stereo Matching in the Twenties

Fabio Tosi, Luca Bartolomei, Matteo Poggi

IJCV|Q1 · IF 11.6|International Journal of Computer Vision

TL;DR

Maps the 2020s of deep stereo: the architectures and paradigms that redefined the field in the last five years, the challenges that stayed open, and a quantitative account of where the benchmarks now stand.

2025
Stereo Anywhere

Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

Luca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano Mattoccia

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Couples a stereo network with monocular priors from a vision foundation model, so the two cover each other's blind spots. Trained only on synthetic data, it still holds up on mirrors, transparencies and textureless regions, where either cue alone fails.

2025
Depth AnyEvent

Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation

Luca Bartolomei, Enrico Mannocci, Fabio Tosi, Matteo Poggi, Stefano Mattoccia

ICCV|A++|International Conference on Computer Vision

TL;DR

Brings monocular depth to event cameras, where dense ground truth essentially does not exist, by distilling a vision foundation model through spatially aligned RGB into dense proxy labels for the event stream.

2025
Active Stereo in the Wild

Active Stereo in the Wild through Virtual Pattern Projection

Luca Bartolomei, Matteo Poggi, Fabio Tosi, Andrea Conti, Stefano Mattoccia

IJCV|Q1 · IF 11.6|International Journal of Computer Vision

TL;DR

Replaces the physical projector of active stereo with a virtual one: sparse measurements from any depth sensor are painted onto both images as patterns consistent with the scene, turning a passive rig into an active one without the projector's range and lighting limits.

2025
CabNIR

CabNIR: A Benchmark for In-Vehicle Infrared Monocular Depth Estimation

Ugo Leone Cavalcanti, Matteo Poggi, Fabio Tosi, Vittorio Cambareri, Vladan Zlokolica, Stefano Mattoccia

WACV|A|Winter Conference on Applications of Computer Vision

TL;DR

A near-infrared benchmark for depth estimation inside the cabin: more than 41,000 frames with ground truth, across 36 vehicles and 45 participants — the scale in-vehicle depth research was missing.

2025
HS-SLAM

HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM

Ziren Gong, Fabio Tosi, Youmin Zhang, Stefano Mattoccia, Matteo Poggi

ICRA|A|International Conference on Robotics and Automation

TL;DR

NeRF-based SLAM with a hybrid encoding — hash grid, tri-planes and one-blob together — plus structural supervision, aimed at the scenes where existing systems lose completeness or drift out of global consistency.

2024
StereoGS

Self-Evolving Depth-Supervised 3D Gaussian Splatting from Rendered Stereo Pairs 🏆 Best Poster Award

Sadra Safadoust, Fabio Tosi, Fatma Güney, Matteo Poggi

BMVC|A|British Machine Vision Conference

TL;DR

The Gaussian Splatting model renders virtual stereo pairs of itself, a stereo network turns them into depth supervision, and that supervision repairs the floating artifacts in its own geometry. A loop that improves as it runs.

2024
Diffusion Models for Monocular Depth Estimation

Diffusion Models for Monocular Depth Estimation: Overcoming Challenging Conditions

Fabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi

ECCV|A++|European Conference on Computer Vision

TL;DR

Generates the hard cases instead of hunting for them: text-to-image diffusion with depth-aware control turns easy scenes into rainy, dark or otherwise adverse ones while preserving their depth, so monocular networks can learn conditions nobody has labelled.

2024
Booster: a Benchmark for Depth from Images of Specular and Transparent Surfaces

Booster: a Benchmark for Depth from Images of Specular and Transparent Surfaces

Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Matteo Poggi, Stefano Mattoccia, Luigi Di Stefano

TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence|vol. 46, no. 1, pp. 85–102

TL;DR

Dense, high-resolution ground truth for exactly what breaks depth estimation — mirrors, glass and other non-Lambertian surfaces — labelled with sub-pixel precision through a deep space-time stereo pipeline.

2024
Neural Disparity Refinement

Neural Disparity Refinement

Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Stefano Mattoccia, Luigi Di Stefano

TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence

TL;DR

The journal version of neural disparity refinement: a continuous formulation that outputs a refined disparity map at any resolution, built for phones, where a high-resolution and a low-resolution camera have to cooperate.

2024
Federated Online Adaptation for Deep Stereo

Federated Online Adaptation for Deep Stereo

Matteo Poggi, Fabio Tosi

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Stereo networks deployed in different environments share what they learn while adapting online. A device that cannot afford to adapt on its own still benefits from the experience of the others.

2023
GO-SLAM

GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction

Youmin Zhang, Fabio Tosi, Stefano Mattoccia, Matteo Poggi

ICCV|A++|International Conference on Computer Vision

TL;DR

Dense neural SLAM that keeps poses and reconstruction globally consistent in real time, through loop closing and online full bundle adjustment, instead of letting tracking error accumulate into a distorted map.

2023
Active Stereo Without Pattern Projector

Active Stereo Without Pattern Projector

Luca Bartolomei, Matteo Poggi, Fabio Tosi, Andrea Conti, Stefano Mattoccia

ICCV|A++|International Conference on Computer Vision

TL;DR

Gives a passive stereo pair the benefits of active stereo with no projector at all: sparse depth hints are virtually projected onto both images as a pattern, so correspondence becomes easy where texture is missing.

2023
NeRF-Supervised Deep Stereo

NeRF-Supervised Deep Stereo

Fabio Tosi, Alessio Tonioni, Daniele De Gregorio, Matteo Poggi

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Trains stereo networks with no ground truth whatsoever. A handheld video becomes a NeRF, the NeRF renders stereo triplets and proxy depth, and the network learns sharp disparities from images that were never captured by a stereo rig.

2023
Learning Depth Estimation for Transparent and Mirror Surfaces

Learning Depth Estimation for Transparent and Mirror Surfaces

Alex Costanzino*, Pierluigi Zama Ramirez*, Matteo Poggi*, Fabio Tosi, Stefano Mattoccia, Luigi Di Stefano

ICCV|A++|International Conference on Computer Vision

TL;DR

Teaches depth for transparent and mirror surfaces without a single annotation: in-paint the offending object, let a monocular model label the repaired image, and fine-tune on those pseudo labels.

2023
GasMono

GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor Scenes

Chaoqiang Zhao, Matteo Poggi, Fabio Tosi, Lingzhe Zhou, Qiyu Sun, Yue Tang, Stefano Mattoccia

ICCV|A++|International Conference on Computer Vision

TL;DR

Self-supervised indoor depth, where large rotations and bare walls break the usual recipe. Coarse poses from multi-view geometry help, but only once the scale ambiguity across scenes is handled — which is the paper's actual contribution.

2022
MonoViT

MonoViT: Self-supervised Monocular Depth Estimation with a Vision Transformer

Chaoqiang Zhao, Youmin Zhang, Matteo Poggi, Fabio Tosi, Xianda Guo, Zheng Zhu, Guan Huang, Yang Tang, Stefano Mattoccia

3DV|A-|International Conference on 3D Vision

TL;DR

Brings the global reasoning of vision transformers to self-supervised monocular depth, where the limited receptive field of convolutions had confined the network to local decisions.

2022
Cross-Spectral Neural Radiance Fields

Cross-Spectral Neural Radiance Fields

Matteo Poggi*, Pierluigi Zama Ramirez*, Fabio Tosi*, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano

3DV|A-|International Conference on 3D Vision

TL;DR

A radiance field shared by cameras that see different parts of the spectrum — colour, multispectral, infrared. It optimizes poses across spectra so that any viewpoint can be rendered in any modality, aligned and at the same resolution.

2022
Booster dataset

Open Challenges in Deep Stereo: the Booster Dataset

Pierluigi Zama Ramirez*, Fabio Tosi*, Matteo Poggi*, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

419 high-resolution indoor samples across 64 scenes, densely annotated and deliberately full of specular and transparent surfaces: a benchmark built around the cases where stereo networks fail.

2022
RGB-Multispectral Matching

RGB-Multispectral Matching: Dataset, Learning Methodology, Evaluation

Fabio Tosi*, Pierluigi Zama Ramirez*, Matteo Poggi*, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Registers colour and multispectral images of very different resolution by treating it as stereo matching, with a new dataset and an architecture trained self-supervised by borrowing a third camera as supervision.

2022
Continual Adaptation for Deep Stereo

Continual Adaptation for Deep Stereo

Matteo Poggi, Alessio Tonioni, Fabio Tosi, Stefano Mattoccia, Luigi Di Stefano

TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence|vol. 44, no. 9, pp. 4713–4729

TL;DR

Adaptation that never stops: rather than assuming the training distribution covers deployment, the stereo network keeps adjusting to whatever environment it actually meets.

2022
On the Confidence of Stereo Matching in a Deep-Learning Era

On the Confidence of Stereo Matching in a Deep-Learning Era: A Quantitative Evaluation

Matteo Poggi, Sunok Kim, Fabio Tosi, Seungryong Kim, Filippo Aleotti, Dongbo Min, Kwanghoon Sohn, Stefano Mattoccia

TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence|vol. 44, no. 9, pp. 5293–5313

TL;DR

A quantitative account of how far confidence estimation for stereo has come in the deep learning era — which measures are worth trusting, and what that reliability buys the algorithms downstream.

2021
Neural Disparity Refinement

Neural Disparity Refinement for Arbitrary Resolution Stereo 🏆 Best Paper Honorable Mention

Filippo Aleotti*, Fabio Tosi*, Pierluigi Zama Ramirez*, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano

3DV|A-|International Conference on 3D Vision

TL;DR

Refines a disparity map at any output resolution through a continuous formulation, so cheap consumer devices — phones with one high-resolution and one low-resolution camera — can produce clean 3D.

2021
Machine learning and binocular stereo survey

On the Synergies Between Machine Learning and Binocular Stereo for Depth Estimation From Images: A Survey

Matteo Poggi, Fabio Tosi, Konstantinos Batsos, Philippos Mordohai, Stefano Mattoccia

TPAMI|Q1 · IF 20.8|IEEE Transactions on Pattern Analysis and Machine Intelligence

TL;DR

Forty years of stereo read through the lens of machine learning: how the two traditions met, where learning genuinely helped, and which problems survived the transition.

2021
SMD-Nets

SMD-Nets: Stereo Mixture Density Networks

Fabio Tosi, Yiyi Liao, Carolin Schmitt, Andreas Geiger

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Predicts a bimodal mixture density instead of a single disparity per pixel. Edges stay sharp where depth jumps, and the output can be sampled at arbitrary resolution.

2020
Distilled Semantics

Distilled Semantics for Comprehensive Scene Understanding from Videos

Fabio Tosi*, Filippo Aleotti*, Pierluigi Zama Ramirez*, Matteo Poggi, Samuele Salti, Stefano Mattoccia, Luigi Di Stefano

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Learns depth, motion and semantics together from monocular video, with the semantic supervision distilled from a pretrained network — no manual labels for any of the three.

2020
Reversing the cycle

Reversing the Cycle: Self-Supervised Deep Stereo through Enhanced Monocular Distillation

Filippo Aleotti*, Fabio Tosi*, Li Zhang, Matteo Poggi, Stefano Mattoccia

ECCV|A++|European Conference on Computer Vision

TL;DR

Reverses the usual direction of self-supervision: instead of stereo teaching a monocular network, a monocular completion network is distilled into a stereo one, which softens the artifacts stereo self-supervision leaves behind.

2019
monoResMatch

Learning Monocular Depth Estimation Infusing Traditional Stereo Knowledge

Fabio Tosi, Filippo Aleotti, Matteo Poggi, Stefano Mattoccia

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

Infers depth from a single image by synthesizing the features of a second view and matching them, importing the machinery of stereo into a monocular network — and training it without any labels, using traditional stereo as the teacher.

2019
Real-time self-adaptive deep stereo

Real-Time Self-Adaptive Deep Stereo Oral

Alessio Tonioni, Fabio Tosi, Matteo Poggi, Stefano Mattoccia

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

A stereo network that fine-tunes itself online while it runs, so accuracy does not collapse the moment the scene stops resembling the training set.

2019
Guided Stereo Matching

Guided Stereo Matching

Matteo Poggi*, Davide Pallotti*, Fabio Tosi, Stefano Mattoccia

CVPR|A++|Conference on Computer Vision and Pattern Recognition

TL;DR

A handful of sparse but reliable depth measurements, fed to the network at inference, steer stereo matching back to accuracy when the environment changes and a fixed model would drift.

2017
Quantitative Evaluation of Confidence Measures

Quantitative Evaluation of Confidence Measures in a Machine Learning World Spotlight

Matteo Poggi, Fabio Tosi, Stefano Mattoccia

ICCV|A++|International Conference on Computer Vision

TL;DR

An extensive comparison of confidence measures for stereo, hand-crafted and learned, across algorithms and datasets — the reference point for telling a good match from a bad one.