3D Object Recognition
Image recognition taught machines to name what is in a photo. 3D recognition teaches them where it is, how big it is and which way it faces, from LiDAR scans, depth cameras and reconstructed scenes.
What 3D object recognition actually is
Read time: 5 min2D image recognition answers one question: what is in this picture? 3D object recognition answers three: what the object is, where it sits in physical space (x, y, z in meters), and how it is oriented. The output is no longer a label or a flat rectangle. It is a labeled 3D box, a labeled set of points, or a full shape.
That extra geometry is why self-driving cars, warehouse robots, AR headsets and drones depend on it. A robot arm cannot grasp a mug from a 2D box. It needs the mug's distance, size and handle direction.
The task family click a task
How 3D data is represented
Read time: 8 minImages have one natural format: a grid of pixels. 3D has several, and the choice of representation drives the choice of model. Point clouds are the closest to raw sensor data and convert easily into the other forms, which is why most modern work starts there.
Lab 1: One object, three representations drag to rotate
Sensors and capture
Read time: 6 minEvery 3D model inherits the strengths and blind spots of its sensor. Knowing the sensor tells you what errors to expect before you train anything.
| Sensor | Range | Density | Weak spot |
|---|---|---|---|
| Spinning / solid-state LiDAR | ~100–250 m | Sparse far away | Rain, fog, black paint |
| RGB-D (structured light, ToF) | ~0.3–5 m | Dense | Sunlight, glass, shiny metal |
| Stereo cameras | Scales with baseline | Dense | Textureless walls, low light |
| Photogrammetry / SfM | Any (offline) | Very dense | Slow, needs many photos |
| Radar (4D imaging) | ~200 m+ | Very sparse | Low resolution, ghost returns |
- LiDAR density falls off with distance, so a car at 60 m may be only a few dozen points.
- Points only exist on surfaces facing the sensor. The back of every object is missing (self-occlusion).
- Sensor fusion (LiDAR + camera) is the production norm: LiDAR for geometry, camera for color and texture.
PointNet and the ordering problem
Read time: 8 minA point cloud is a set. The same chair can arrive as points in any order, and the answer must not change. A normal neural network, which reads inputs in a fixed order, breaks this rule. Before 2017 researchers dodged the issue by converting points to voxels or images, losing detail and burning memory.
PointNet (Qi et al., CVPR 2017) solved it directly. It runs the same small network (a shared MLP) on every point independently, then combines all point features with a symmetric function: max pooling. Max does not care about order, so the global feature is identical however the points are shuffled. A small learned transform (T-Net) also aligns the input to reduce sensitivity to rotation.
PointNet++ added local neighborhoods: sample centers, group nearby points, apply PointNet to each group, repeat. That gives the hierarchy of local-to-global features CNNs get from stacked convolutions.
Lab 2: Why max pooling is order-proof shuffle the points
Naive: concatenate raw points
PointNet: shared MLP + max pool
3D object detection and IoU
Read time: 8 minDetection finds every object in a whole scene, which can hold over 100,000 LiDAR points. Three families dominate:
- Voxel-based (VoxelNet, SECOND): voxelize, apply sparse 3D convolutions, predict boxes. Accurate, heavier.
- Pillar-based (PointPillars): group points into vertical columns, encode each with a mini PointNet, flatten to a bird's-eye-view (BEV) pseudo-image, then run fast 2D convolutions. The original paper reported 62 Hz on KITTI.
- Center-based and fusion (CenterPoint, BEVFusion, TransFusion): predict object centers on a BEV map and fuse camera features. Common on nuScenes and Waymo leaderboards.
Predictions are scored by Intersection over Union (IoU): the overlap volume divided by the combined volume. On KITTI, a predicted car only counts if its 3D IoU with the ground truth is at least 0.7. Small errors in height or position punish 3D IoU much more than 2D IoU.
Lab 3: 3D IoU calculator car box: 4.5 × 2.0 × 1.6 m
Transformers and 3D foundation models
Read time: 6 minSelf-attention is naturally permutation-invariant, so transformers fit point sets well. Point Transformer (V1 to V3) became the leading backbone for 3D segmentation. Point Transformer V3 (CVPR 2024 oral) traded complex neighbor search for serialization: it orders points along space-filling curves so attention runs fast on large scenes.
The 2025 shift is self-supervised pretraining, the same move DINO made for images. Meta's Sonata (CVPR 2025 Highlight) pretrains PTv3 on about 140,000 unlabeled scenes. Its features are strong enough that a simple linear probe reaches 72.5% mIoU on ScanNet segmentation. Follow-up work such as Concerto distills 2D image foundation models into the 3D encoder.
A parallel trend is open-vocabulary 3D: lifting CLIP-style language features onto points so you can query a scan for "fire extinguisher" without training a class for it.
NeRF and Gaussian Splatting
Read time: 6 min (plus optional videos)Sometimes there is no LiDAR, only photos. Two techniques turn photos into 3D scenes that recognition can then run on.
NeRF (Neural Radiance Fields) trains a network to return color and density for any 3D point and viewing direction. It renders by ray marching, which is high quality but slow.
3D Gaussian Splatting (Kerbl et al., 2023) represents the scene as millions of small 3D Gaussian blobs with position, shape, opacity and view-dependent color. They are rasterized, not ray-traced, so scenes render in real time. Recognition research now attaches semantic features to each Gaussian, giving segmentable, queryable scenes.
Best practices and pitfalls
Read time: 7 minBest practices
- Pick the representation for the job: points for precision, voxels or pillars for speed, BEV for driving.
- Normalize each sample: center it, scale to a unit sphere, and fix a point count (PointNet used 1,024 for classification).
- Augment in 3D: random rotation around the vertical axis, scaling, jitter, point dropout, and "copy-paste" of objects between scenes.
- Report results by distance bin and occlusion level, not one average. Far objects fail first.
- Start from pretrained backbones (PTv3/Sonata, or OpenPCDet and MMDetection3D model zoos) before training from scratch.
- Visualize predictions in 3D (Open3D, CloudCompare, rerun) every iteration. Metrics hide flipped headings.
Lab 4: Spot the pitfall sound practice or mistake?
Privacy, safety and Ask Me Anything
Read time: 6 min3D scans capture more than objects. Home and office scans record floor plans, belongings and sometimes people. LiDAR street data can reveal license plates through camera fusion and identify individuals by gait or body shape. Treat scan data as personal data where GDPR or similar laws apply.
- Blur or remove faces and plates in fused camera data; strip or downsample people in shared point clouds.
- Document dataset geography and weather. Models trained on sunny-climate data fail in snow.
- In safety-critical use (vehicles, robots near people), test long-tail cases: children, wheelchairs, unusual vehicles.
- Check dataset licenses. Several popular driving datasets are non-commercial only.
AMA: questions people usually ask
Is 3D recognition just 2D recognition with an extra axis?
No. 3D data is sparse, unordered and unevenly sampled, so standard convolutions do not apply directly. That is why point-specific architectures like PointNet, sparse convolutions and point transformers exist.
Can I do 3D recognition with only a phone camera?
Yes, with limits. Monocular 3D detection and depth estimation work from single images but are less accurate at estimating distance. iPhones and iPads with LiDAR, or photo-to-splat apps, give much better geometry.
Which library should I start with?
Open3D for loading and visualizing point clouds, PyTorch3D or Pointcept for research models, OpenPCDet or MMDetection3D for LiDAR detection, and MATLAB's Lidar Toolbox if your team is already in MATLAB.
What datasets should I know?
ModelNet40 and ShapeNet (object classification), ScanNet and S3DIS (indoor segmentation), KITTI, nuScenes and Waymo Open (driving detection).
Do I need a GPU?
For training, yes. For inference, small models like PointPillars run on embedded GPUs such as NVIDIA Jetson, and classification of single objects runs on a laptop CPU.
Where does this show up outside self-driving?
Warehouse picking robots, construction progress tracking against BIM models, retail shelf auditing, AR furniture placement, surgical navigation, and drone surveys for agriculture and insurance.
How does this connect to digital humans or avatars?
Body and face reconstruction uses the same tools: depth sensing, mesh fitting (SMPL body models) and increasingly Gaussian avatars, which are animated splat-based people rendered in real time.
What should I learn next?
Run a pretrained PointNet++ on ModelNet40, then a PointPillars model on a KITTI sample, then capture your own scene with a phone and turn it into a Gaussian splat.
Test yourself: 10 questions
Glossary
14 terms
- Point cloud
- An unordered set of 3D points, often with intensity or color.
- Voxel
- A cell in a regular 3D grid; the 3D equivalent of a pixel.
- Pillar
- A vertical column of space used by PointPillars to group LiDAR points.
- BEV
- Bird's-eye view: a top-down 2D projection of a 3D scene.
- LiDAR
- Light Detection and Ranging; measures distance with laser pulses.
- RGB-D
- A color image paired with a per-pixel depth map.
- Permutation invariance
- Output does not change when input order changes.
- Symmetric function
- A function such as max or sum whose result ignores argument order.
- T-Net
- PointNet's small network that predicts an alignment transform.
- 3D IoU
- Overlap volume divided by union volume of two 3D boxes.
- 6-DoF pose
- Three position plus three rotation values describing an object's placement.
- Sparse convolution
- Convolution computed only on occupied voxels.
- NeRF
- Neural Radiance Field: a network mapping 3D position and view to color and density.
- Gaussian Splatting
- Scene representation of many 3D Gaussians, rasterized in real time.
Sources
- Qi et al., PointNet: Deep Learning on Point Sets (CVPR 2017). arxiv.org/pdf/1612.00593
- Lang et al., PointPillars: Fast Encoders for Object Detection from Point Clouds. arxiv.org/pdf/1812.05784
- Lu et al., Transformers in 3D Point Clouds: A Survey. arxiv.org/abs/2205.07417
- Zhu et al., Transformer-Based 3D Object Detection for Autonomous Driving survey, Drones (2024). mdpi.com
- Wu et al., Sonata: Self-Supervised Learning of Reliable Point Representations (CVPR 2025). openaccess.thecvf.com
- Sonata code and checkpoints. github.com/facebookresearch/sonata
- Project Aria, Introducing Sonata. projectaria.com
- 3D Gaussian as a New Vision Era: A Survey. arxiv.org/abs/2402.07181
- MathWorks, Lidar 3-D Object Detection Using PointPillars. mathworks.com
- MathWorks, Point Cloud Classification Using PointNet. mathworks.com
- Carós, PointNet Explained Visually. medium.com
- Stanford CS231n schedule (Lecture 15: 3D Vision). cs231n.stanford.edu
- Georgia Tech CS 6476 lecture slides on PointNet. faculty.cc.gatech.edu