Elective in Artificial Intelligence · 3 CFU module

Foundations of Spatial AI

How machines perceive, reconstruct, and render the 3D world: from classical geometry and image formation, through depth estimation, Structure-from-Motion and SLAM, to modern neural rendering with NeRF and 3D Gaussian Splatting.

Credits
3 CFU
Hours
~30 h
Level
MSc AI and Robotics
Language
English
Semester
A.Y. 2026/27First period
23 Sep – 4 Nov 2026
About this module

This module follows a single thread: turning 2D observations into 3D structure you can reconstruct from real images and render back into new views. We begin with how images are formed and the classical ways of representing 3D geometry, move through depth estimation and multi-view reconstruction and state estimation (Structure-from-Motion and SLAM), and then into the neural rendering methods: volumetric fields, NeRF, and 3D Gaussian Splatting, that now define the field.

Teaching combines lectures with hands-on practicals that run inside the class hours, on a laptop, using Python and PyTorch with COLMAP and NeRFStudio. The module is assessed through a short individual quiz and a group project or group presentation. The last two sessions are seminars that look past the exam program, at where this field is heading. It is one module of the 12 CFU Elective in Artificial Intelligence.

Learning outcomes
  • Represent 3D scenes with point clouds, meshes, voxels, implicit fields, and Gaussians, and reason about the geometry of image formation.
  • Reconstruct geometry from images: monocular and multi-view depth, Structure-from-Motion, and SLAM.
  • Build and train neural rendering pipelines: volume rendering, NeRF, and 3D Gaussian Splatting, and weigh their trade-offs.
  • Diagnose what breaks in a 3D pipeline: pose and intrinsics conventions, captures that defeat reconstruction, and sampling error in volume rendering.

What you will learn

A glimpse of what you build during the module.

Mesh vs Rendering

The same office scene reconstructed two ways: a classical mesh and rendering (3D Gaussian Splatting). Drag the slider to wipe between them.

MrHash

Preliminaries

This module follows a single chain from end to end: how an image is formed, how a set of images becomes 3D geometry, and how that geometry is rendered back into new views. The program stops at rendering, which leaves each step the time it needs. By the end you should understand why each link in that chain is built the way it is.

Treat the side list as a comfort check. You do not need to arrive knowing everything: it is enough to recognize these ideas and not freeze when they appear on a slide. Everything past that we cover together, and we explain what we use as we use it.

Comfort check
  • Linear algebra: matrices, least squares, SVD
  • Rigid transforms: rotation matrices, homogeneous coordinates
  • Calculus: gradients and the chain rule
  • Probability: Gaussians, mean and covariance
  • Python with NumPy
  • PyTorch: tensors, autograd, a basic training loop
  • Neural network basics: CNNs, loss functions, backpropagation

Epipolar geometry, multi-view stereo, Structure-from-Motion and SLAM, volume rendering, NeRF, and Gaussian Splatting are all introduced from scratch during the module. No prior exposure is assumed.

You will not be examined on linear algebra. With thirteen sessions on one chain, we can go deep enough for you to build something and understand why it works. Several of these topics are entire courses, or entire PhDs, on their own. If one of them grabs you, the suggested readings in the schedule are the door in (this will be updated during the course), and we're happy to point you further at office hours.

Logistics

Teaching sessions
2 per week
Mon 08:00–10:00
Wed 08:00–11:00
Room
Room B2
DIAG · Via Ariosto 25
Period
23 Sep – 4 Nov 2026
Six weeks · 13 sessions
Office hours
By appointment
Day · time TBD
Enrollment form

Students taking this module are strongly encouraged to fill in the enrollment form. It is how we reach you about slides, practicals, room changes and anything else, so it makes communication with the teachers much easier.

Fill in the form ↗

Teaching staff

Emanuele Giacomini
Emanuele Giacomini

Schedule

Thirteen sessions over six weeks, twice a week: Monday 08:00–10:00 and Wednesday 08:00–11:00 in room B2. Lessons start on Wednesday 23 September and end on Wednesday 4 November 2026. Each row below is a block of the course, so the order of topics inside a block can shift as we go. Slides are shared through Google Classroom with enrolled students only. Readings are added under each topic during the lessons. SCHEDULE CAN BE SLIGHTLY CHANGED.

Block
Dates
What we cover
Block 1 Foundations 2 sessions
23 – 28 Sep
Course overview
Image formation
3D representations
Block 2 Reconstruction 4 sessions
30 Sep – 12 Oct
Single-view depth and LiDAR completion
Depth Map Prediction from a Single Image using a Multi-Scale Deep Network, Eigen et al. (2014)
Unsupervised Monocular Depth Estimation with Left-Right Consistency, Godard et al. (2017)
Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer (MiDaS), Ranftl et al. (2020)
Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation (Marigold), Ke et al. (2024)
Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion, Viola et al. (2024)
Multi-view depth estimation and MVS
Structure-from-Motion
SLAM
Block 3 Neural rendering I 3 sessions
14 – 21 Oct
Volume rendering
Neural Radiance Fields (NeRF)
Neural implicit rendering
Block 4 Neural rendering II 2 sessions
26 – 28 Oct
Differentiable primitive rendering
Neural appearance modeling
3D Gaussian Splatting
Seminars · 2 – 4 Nov
Block 5 Seminars 2 sessions
2 – 4 Nov
Seminar session 1
Seminar session 2

Grading & exam

Individual quiz · 13 questions10 pts
Group project or presentation20 pts +3 bonus
Two small homeworks+3 bonus
Exam format

Assessment has two main parts: an individual quiz (13 questions, max 10 points), and either a group project or a group presentation (max 20 points). Both are done in groups of two or three students. Two small homeworks during the course are worth up to +3 points, and a further +3 go to outstanding projects and outstanding presentations.

This grade counts toward the overall mark of the 12 CFU Elective in Artificial Intelligence.