CLIKT Lessons
3D Computer Vision · Self-paced lesson

3D Object Recognition

Image recognition taught machines to name what is in a photo. 3D recognition teaches them where it is, how big it is and which way it faces, from LiDAR scans, depth cameras and reconstructed scenes.

60 minTime
IntermediateLevel
9 modulesFormat
4 labs + quizHands-on
Module 1 · Foundations

What 3D object recognition actually is

Read time: 5 min

2D image recognition answers one question: what is in this picture? 3D object recognition answers three: what the object is, where it sits in physical space (x, y, z in meters), and how it is oriented. The output is no longer a label or a flat rectangle. It is a labeled 3D box, a labeled set of points, or a full shape.

That extra geometry is why self-driving cars, warehouse robots, AR headsets and drones depend on it. A robot arm cannot grasp a mug from a 2D box. It needs the mug's distance, size and handle direction.

The task family click a task

2D vs 3D in one line. A 2D detector outputs 4 numbers per object (a box in pixels). A typical 3D detector outputs 7: center x, y, z, length, width, height, and heading angle, all in meters and radians.
Module 2 · Data

How 3D data is represented

Read time: 8 min

Images have one natural format: a grid of pixels. 3D has several, and the choice of representation drives the choice of model. Point clouds are the closest to raw sensor data and convert easily into the other forms, which is why most modern work starts there.

Point cloudUnordered list of (x, y, z) points, plus intensity or color. Sparse and exact.
Voxels3D pixels on a grid. Easy for 3D CNNs, but mostly empty and memory-hungry.
MeshVertices and triangles. Great for graphics and CAD, harder for learning.
RGB-D / depthA normal image plus a distance per pixel. Often called 2.5D.
Multi-viewMany 2D renders or photos of one object, fused by a 2D network.
Implicit / splatsA neural field (NeRF) or 3D Gaussians describing the whole scene.

Lab 1: One object, three representations drag to rotate

Try this. Switch to Voxel grid and drop resolution to 6. The torus hole disappears: that is quantization error. Then raise it to 32 and watch the occupancy figure. Over 90% of the grid is empty, which is why sparse convolutions exist.
Module 3 · Capture

Sensors and capture

Read time: 6 min

Every 3D model inherits the strengths and blind spots of its sensor. Knowing the sensor tells you what errors to expect before you train anything.

SensorRangeDensityWeak spot
Spinning / solid-state LiDAR~100–250 mSparse far awayRain, fog, black paint
RGB-D (structured light, ToF)~0.3–5 mDenseSunlight, glass, shiny metal
Stereo camerasScales with baselineDenseTextureless walls, low light
Photogrammetry / SfMAny (offline)Very denseSlow, needs many photos
Radar (4D imaging)~200 m+Very sparseLow resolution, ghost returns
Module 4 · Core model

PointNet and the ordering problem

Read time: 8 min

A point cloud is a set. The same chair can arrive as points in any order, and the answer must not change. A normal neural network, which reads inputs in a fixed order, breaks this rule. Before 2017 researchers dodged the issue by converting points to voxels or images, losing detail and burning memory.

PointNet (Qi et al., CVPR 2017) solved it directly. It runs the same small network (a shared MLP) on every point independently, then combines all point features with a symmetric function: max pooling. Max does not care about order, so the global feature is identical however the points are shuffled. A small learned transform (T-Net) also aligns the input to reduce sensitivity to rotation.

PointNet++ added local neighborhoods: sample centers, group nearby points, apply PointNet to each group, repeat. That gives the hierarchy of local-to-global features CNNs get from stacked convolutions.

Lab 2: Why max pooling is order-proof shuffle the points

Naive: concatenate raw points

PointNet: shared MLP + max pool

PointNet's catch. Max pooling keeps a small set of "critical points" that define the shape, which makes it robust to missing points. But plain PointNet sees no local structure, so it struggles with fine detail. That is the gap PointNet++, DGCNN (dynamic graph edges) and later transformers close.
Module 5 · Detection

3D object detection and IoU

Read time: 8 min

Detection finds every object in a whole scene, which can hold over 100,000 LiDAR points. Three families dominate:

Predictions are scored by Intersection over Union (IoU): the overlap volume divided by the combined volume. On KITTI, a predicted car only counts if its 3D IoU with the ground truth is at least 0.7. Small errors in height or position punish 3D IoU much more than 2D IoU.

Lab 3: 3D IoU calculator car box: 4.5 × 2.0 × 1.6 m

Notice. A 0.4 m forward error plus 0.2 m in height already costs a lot of 3D IoU. That is why 3D benchmarks also report BEV IoU and why nuScenes uses a center-distance metric (NDS) instead of pure IoU.
Module 6 · State of the art

Transformers and 3D foundation models

Read time: 6 min

Self-attention is naturally permutation-invariant, so transformers fit point sets well. Point Transformer (V1 to V3) became the leading backbone for 3D segmentation. Point Transformer V3 (CVPR 2024 oral) traded complex neighbor search for serialization: it orders points along space-filling curves so attention runs fast on large scenes.

The 2025 shift is self-supervised pretraining, the same move DINO made for images. Meta's Sonata (CVPR 2025 Highlight) pretrains PTv3 on about 140,000 unlabeled scenes. Its features are strong enough that a simple linear probe reaches 72.5% mIoU on ScanNet segmentation. Follow-up work such as Concerto distills 2D image foundation models into the 3D encoder.

A parallel trend is open-vocabulary 3D: lifting CLIP-style language features onto points so you can query a scan for "fire extinguisher" without training a class for it.

Practical takeaway. In 2026 you rarely train a point backbone from scratch. Start from a pretrained PTv3/Sonata checkpoint (open-sourced via the Pointcept codebase) and fine-tune or linear-probe on your data.
Module 7 · Reconstruction

NeRF and Gaussian Splatting

Read time: 6 min (plus optional videos)

Sometimes there is no LiDAR, only photos. Two techniques turn photos into 3D scenes that recognition can then run on.

NeRF (Neural Radiance Fields) trains a network to return color and density for any 3D point and viewing direction. It renders by ray marching, which is high quality but slow.

3D Gaussian Splatting (Kerbl et al., 2023) represents the scene as millions of small 3D Gaussian blobs with position, shape, opacity and view-dependent color. They are rasterized, not ray-traced, so scenes render in real time. Recognition research now attaches semantic features to each Gaussian, giving segmentable, queryable scenes.

Want the full lecture? Stanford CS231n's 3D Vision lecture covers representations, PointNet and mesh prediction. Prior-year recordings are on YouTube; the schedule is in Sources.
Module 8 · Practice

Best practices and pitfalls

Read time: 7 min

Best practices

Lab 4: Spot the pitfall sound practice or mistake?

Module 9 · Responsibility + AMA

Privacy, safety and Ask Me Anything

Read time: 6 min

3D scans capture more than objects. Home and office scans record floor plans, belongings and sometimes people. LiDAR street data can reveal license plates through camera fusion and identify individuals by gait or body shape. Treat scan data as personal data where GDPR or similar laws apply.

AMA: questions people usually ask

Is 3D recognition just 2D recognition with an extra axis?

No. 3D data is sparse, unordered and unevenly sampled, so standard convolutions do not apply directly. That is why point-specific architectures like PointNet, sparse convolutions and point transformers exist.

Can I do 3D recognition with only a phone camera?

Yes, with limits. Monocular 3D detection and depth estimation work from single images but are less accurate at estimating distance. iPhones and iPads with LiDAR, or photo-to-splat apps, give much better geometry.

Which library should I start with?

Open3D for loading and visualizing point clouds, PyTorch3D or Pointcept for research models, OpenPCDet or MMDetection3D for LiDAR detection, and MATLAB's Lidar Toolbox if your team is already in MATLAB.

What datasets should I know?

ModelNet40 and ShapeNet (object classification), ScanNet and S3DIS (indoor segmentation), KITTI, nuScenes and Waymo Open (driving detection).

Do I need a GPU?

For training, yes. For inference, small models like PointPillars run on embedded GPUs such as NVIDIA Jetson, and classification of single objects runs on a laptop CPU.

Where does this show up outside self-driving?

Warehouse picking robots, construction progress tracking against BIM models, retail shelf auditing, AR furniture placement, surgical navigation, and drone surveys for agriculture and insurance.

How does this connect to digital humans or avatars?

Body and face reconstruction uses the same tools: depth sensing, mesh fitting (SMPL body models) and increasingly Gaussian avatars, which are animated splat-based people rendered in real time.

What should I learn next?

Run a pretrained PointNet++ on ModelNet40, then a PointPillars model on a KITTI sample, then capture your own scene with a phone and turn it into a Gaussian splat.

Knowledge check

Test yourself: 10 questions

Reference

Glossary

14 terms
Point cloud
An unordered set of 3D points, often with intensity or color.
Voxel
A cell in a regular 3D grid; the 3D equivalent of a pixel.
Pillar
A vertical column of space used by PointPillars to group LiDAR points.
BEV
Bird's-eye view: a top-down 2D projection of a 3D scene.
LiDAR
Light Detection and Ranging; measures distance with laser pulses.
RGB-D
A color image paired with a per-pixel depth map.
Permutation invariance
Output does not change when input order changes.
Symmetric function
A function such as max or sum whose result ignores argument order.
T-Net
PointNet's small network that predicts an alignment transform.
3D IoU
Overlap volume divided by union volume of two 3D boxes.
6-DoF pose
Three position plus three rotation values describing an object's placement.
Sparse convolution
Convolution computed only on occupied voxels.
NeRF
Neural Radiance Field: a network mapping 3D position and view to color and density.
Gaussian Splatting
Scene representation of many 3D Gaussians, rasterized in real time.
Reference

Sources

  1. Qi et al., PointNet: Deep Learning on Point Sets (CVPR 2017). arxiv.org/pdf/1612.00593
  2. Lang et al., PointPillars: Fast Encoders for Object Detection from Point Clouds. arxiv.org/pdf/1812.05784
  3. Lu et al., Transformers in 3D Point Clouds: A Survey. arxiv.org/abs/2205.07417
  4. Zhu et al., Transformer-Based 3D Object Detection for Autonomous Driving survey, Drones (2024). mdpi.com
  5. Wu et al., Sonata: Self-Supervised Learning of Reliable Point Representations (CVPR 2025). openaccess.thecvf.com
  6. Sonata code and checkpoints. github.com/facebookresearch/sonata
  7. Project Aria, Introducing Sonata. projectaria.com
  8. 3D Gaussian as a New Vision Era: A Survey. arxiv.org/abs/2402.07181
  9. MathWorks, Lidar 3-D Object Detection Using PointPillars. mathworks.com
  10. MathWorks, Point Cloud Classification Using PointNet. mathworks.com
  11. Carós, PointNet Explained Visually. medium.com
  12. Stanford CS231n schedule (Lecture 15: 3D Vision). cs231n.stanford.edu
  13. Georgia Tech CS 6476 lecture slides on PointNet. faculty.cc.gatech.edu