Peder Bergebakken Sundt
Theoharis Theoharis
Differentiable rendering requires good gradient flow
Primitives
Differentiable rendering requires good gradient flow
Neural Fields
Approximate representations?
Primitives
Neural Fields
why pick one over the other?
Approximate representations?
Differentiable rendering requires good gradient flow
Primitives
why pick one over the other?
Neural Ray Fields
Neural Fields
Differentiable rendering requires good gradient flow
Primitives
Neural Fields
Why pick one over the other?
The current SOTA
feature great render times.
Minimize primitive processing
Minimize network evaluations
Image Diffusion
3D Gaussian Splatting
Voxel Grids
Fast Dipole Sums
Implicit Surface / Radiance Fields
Ray Intersection / Light Fields
Ratio tracking
Edge-sampling
Rasterization
Ray Marching
Ray Casting
Neural Fields
Primitives
Speedup
(closed-form solution)
(iterative solution)
Neural Ray fields are compact and feature non-diverging memory access,
making them attractive as intersection
shaders in hybrid path-tracing
Figure:
The Vulkan Ray-Tracing Pipeline
A simplified timeline
PRIF
MARF
PMARF
PDDF
LFN
Light Fields
Intersection Fields
NFD
AutoInt
5D
4D
5D
4D
𝒩-BVH
LSNIF
Hybrid
SRDF
Hybrid
A simplified timeline
Orthogonal to our work.
Plücker
embeddings
Plücker
embeddings
Hybrid
PRIF
MARF
PMARF
PDDF
LFN
This Work
Light Fields
Intersection Fields
NFD
AutoInt
5D
4D
5D
4D
𝒩-BVH
LSNIF
Hybrid
SRDF
Hybrid
A simplified timeline
We iterate on MARF and PMARF, with some improvements
inspired by NFD and PDDF, and some novel.
The 5 (or 4) degrees-of-freedom in rays
permit view-dependent geometry.
You can alleviate view overfitting with
dense multi-view supervision.
...but can one do without?
Early work
⇒ Multi-view consistent priors
and regularization
Overfitting on purpose is a no-go.
shape
Ray Marching?
shape
Lipschitz continuity
k inscribed spheres
"medial atoms"
depth
"Topological Skeleton"
Medial Axis / Surface
Medial atoms are easy to regularize, and provide good multi-view priors
k inscribed spheres
"medial atoms"
"Topological Skeleton"
Medial Axis / Surface
Medial atoms are easy to regularize, and provide good priors / inductive bias
k inscribed spheres
"medial atoms"
"Topological Skeleton"
Medial Axis / Surface
Medial atoms are easy to regularize, and provide good priors / inductive bias
Number of candidate predictions per ray.
... but increasing k beyond 16 does not improve reconstruction quality much in practice...
Why?
k determines:
They works great for color rendering,
but not for solid surfaces.
Gaussians (or metaballs) are...
It works great for color
Consider the ray-to-surface mapping
Consider the ray-to-surface mapping
Smoothness
through erosion.
Silhouettes unaffected.
View dependent.
Figure:
Softmin atom normals
Atoms cooperate and specialize to
different parts of the shape
Atoms compete to fit the shape
Train:
Test:
Train:
Test:
Train:
Test:
We start training at (c) then decay to (f)
Controls blending sharpness/entropy
First some preliminaries on parametric MARFs
A ray
k atoms
Method
A ray
k atoms
k coordinates
Method
In essence, fit k
parametric surfaces
Couples the atom center and radius,
forms a consistent medial surface.
Shape is however limited
to k medial planes.
A ray
k atoms
k coordinates
Method
Requires a parameter-domain jump w.r.t input ray
In essence, fit k
parametric surfaces
A ray
k atoms
k coordinates
Method
The 3D unit sphere
The 3D unit sphere
This works for any axis, provided no cycles are formed.
(i.e. on genus-0 subsets)
The 3D unit sphere
The union of two such
shapes may have any genus.
Equation: Objective Function
Constrain to unit sphere
Maximize utilization
The 3D unit sphere
Equation: Objective Function
Required to leverage
wrap-around sphere topology
The 3D unit sphere
Constrain to unit sphere
Maximize utilization
Dense Data
Data Synthesis
Data Augmentation
With a coarse-to-fine training schedule,
to not compromise the final fidelity.
Gaussian Beam Distribution
Slope
Inspired by Rebain et al. (2024) we permute rays with a
Rebain D, Yazdani S, Yi KM, Tagliasacchi A. Neural fields as distributions: Signal processing beyond Euclidean space. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 Jun 16 (pp. 4274-4283). IEEE.
Slope
Waist
Gaussian Beam Distribution
Inspired by Rebain et al. (2024) we permute rays with a
Rebain D, Yazdani S, Yi KM, Tagliasacchi A. Neural fields as distributions: Signal processing beyond Euclidean space. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 Jun 16 (pp. 4274-4283). IEEE.
Slope
Waist
Rebain D, Yazdani S, Yi KM, Tagliasacchi A. Neural fields as distributions: Signal processing beyond Euclidean space. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 Jun 16 (pp. 4274-4283). IEEE.
Gaussian Beam Distribution
We decay from a wide beam to
a tight beam during training.
Inspired by Rebain et al. (2024) we permute rays with a
We permute the predicted
Parameter Domain Coordinates
with decaying Gaussian noise
Works better when we don't
clamp the radial component
We permute the predicted
Parameter Domain Coordinates
We decay the softmin blending temperature T
(sharpness/entropy)
Soft/"see through",
good gradient flow
Sharp/high fidelity,
reduced gradient flow
We decay the softmin blending temperature T
(sharpness/entropy)
We decay the amount of dropout from excessive to minimal,
in effect band-limiting the initial fit to avoid local minima.
I.e. how the gradient is allocated between the predicted atoms.
With softmin the gradient propagate to
multiple atoms, but only true hits.
Furthermore, the first hit still receive
the majority of the signal.
Furthermore, the first hit still receive
the majority of the signal.
With softmin the gradient propagate to
multiple atoms, but only true hits.
MARF blends atoms
using the argmin of this metric
Miss predictions are discarded
Prefers the hit furthest along negative ray direction
Invariant of hit/miss predictions,
only the ground truth matters
This metric however requires ground truths, making it
only apply to training, not to validation or testing.
We instead prefer the atom that
has to move the least
⇒ Repair false misses,
improving recall.
⇒ View-invariant
softmin blending.
Figure:
Reconstruction error of
a training view and
a novel view, 15° apart.
Shading: Chamfer distance
to ground truth mesh.
Our method repairs the silhouette missed by MARF
The MARF loss is good for surface details,
but struggles with atom specialization,
false misses/poor recall, and has
a hit/miss imbalance.
Atom Displacement Loss
Atom Displacement loss
Euclidean Normal Loss
Early MV loss
Truncated regularization
Parameter domain
regularization
Can pull in opposite directions!
Applies to true hits only
Pushes atom along ray
to fit target depth
Pivots atom about hit
to fit target surface normal
Recall our SAS:
Prefers the atom that
has to move the least.
What if we penalize
this distance?
Applies to all hits
Supervises intersection points
and normals jointly
Cosine Distance
Euclidean Distance
Yin R, Chen Y, Karaoglu S, Gevers T. Ray-Distance Volume Rendering for Neural Scene Reconstruction. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G, editors. Computer Vision – ECCV 2024, Cham: Springer Nature Switzerland; 2025, p. 377–94.
https://doi.org/10.1007/978-3-031-72630-9_22
Smooth and stable
Sharp, but unstable
First, a small recap
Double backpropagation
... how does differentiating w.r.t. the ray direction help?
Ray
Hit
Ray
Hit
4D embedding
Global along-ray translation invariance
by construction through normalization
or
Ray origin is normalized
before network ever sees it,
but is still a part of the
auto-differentiation graph!
Ray
Hit
4D embedding
or
penalizing changes w.r.t. viewpoint.
Global along-ray translation invariance
by construction through normalization
Ray origin is normalized
before network ever sees it,
but is still a part of the
auto-differentiation graph!
Fix intersected atom center and radius, not just the intersection point.
Ray Decoder
Wastes capacity to accommodate
a multi-view inconsistent ray decoder
Medial Surface Decoders
Weighted with softmin weights
A brief recap
adds a constant positive pressure on atom radii (MAT maximality).
limits per-candidate area of influence, to specialize atoms to separate parts and handle discontinuities.
there is a single trivial solution
when no other loss apply.
Failure mode:
adds a constant positive pressure on atom radii (MAT maximality).
limits per-candidate area of influence, to specialize atoms to separate parts and handle discontinuities.
there is a single trivial solution
when no other loss apply.
Failure mode:
The majority of predictions miss,
from a per-candidate perspective.
adds a constant positive pressure on atom radii (MAT maximality).
limits per-candidate area of influence, to specialize atoms to separate parts and handle discontinuities.
there is a single trivial solution
when no other loss apply.
Failure mode:
The majority of predictions miss,
from a per-candidate perspective.
3 architectures
2 setups
( + PRIF )
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Depth-based
(non-medial)
Medial baselines
Baseline
with our S²
Our softmin & improved loss
Chamfer Distance
Cosine Similarity
Medial Atom Normals
Differential Normals
Intersection over Union
Precision
Recall
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Table: Reconstruction scores. Average of 19 shapes.
Same correctness, with improved capacity and recall.
Best overall.
Start to fall behind.
Has a "indecisive"
fit from the higher gradient flow.
fails to generalize
multi-view stable
multi-view stable
multi-view stable
multi-view stable
multi-view stable
degeneracies near
overhangs
degeneracies near
overhangs
limited topology
limited topology
fails to generalize
multi-view stable
multi-view stable
multi-view stable
multi-view stable
multi-view stable
degeneracies near
overhangs
degeneracies near
overhangs
missing silhouette
missing silhouette
missing silhouette
limited topology
limited topology
fails to generalize
multi-view stable
multi-view stable
multi-view stable
multi-view stable
multi-view stable
degeneracies near
overhangs
degeneracies near
overhangs
missing silhouette
missing silhouette
missing silhouette
limited topology
limited topology
problems with inorganic shapes
that feature sharp angles
Specialized
Specialized
Entangled atoms
Limited topology
Vanishing radius
Specialized
Specialized
Limited topology
Vanishing radius
A novel-view problem,
training views are unaffected.
Likely a ray-decoder interpolation issue.
This hurdle may be fixed with:
Rebain D, Yazdani S, Yi KM, Tagliasacchi A. Neural fields as distributions: Signal processing beyond Euclidean space. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 Jun 16 (pp. 4274-4283). IEEE.
Entangled atoms
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Sundt PB, Theoharis T. Towards multi-view consistency in neural ray fields using parametric medial surfaces. Computers & Graphics 2024;123:103991.
k-means init
Agnostic init
Data
Table: Ablation studies
Table: Ablation studies
but is benefitial on PMARF baselines.
Reinforces characteristics observed earlier.
Detrimental on Our R² configuration,
Table: Ablation studies
Table: Ablation studies
argmin
softmin
data
Table: Ablation studies
Table: Ablation studies
⇒ Our Atom Displacement Loss requires Supervised Atom Selection.
our SAS mitigates the issue.
Penalizing normals early is unstable (MARF baseline also avoids it),
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Strictly improved
Detrimental?
Primarily on mechanical shapes
Table: Ablation studies
Table: Ablation studies
Ray augmentation improves fidelity and recall.
And so does parameter domain augmentation.
Table: Ablation studies
Table: Ablation studies
And so does parameter domain augmentation.
Ray augmentation improves fidelity and recall.
Table: Ablation studies
Like MARF baseline
Hits only
Table: Ablation studies
Our 0.3 threshold strikes a sweetspot between regularizing all rays or regularizing hits only
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Table: Ablation studies
Excessive
Insufficient
Sweetspot?
Decaying dropout
Fixed dropout
Our decay from excessive to minimal avoids
local minima without sacrificing fidelity,
but rather improve it.
Table: Ablation studies
Cosine and Euclidean
Euclidean-only
Cosine-only
Stable, but smooth
Detailed, but unstable
A balance
Table: Timing Results
Ours have more losses, more gradients.
S² has k more activations and a vector normalization.
Table: Timing Results
Performance is reclaimed by our early multi-view loss.
Ours have more losses, more gradients.
S² has k more activations and a vector normalization.
Table: Timing Results
Parametric networks excel when compute bound.
Softmin blending is less divergent than Argmin lookups.
Table: Timing Results
Large but simple MLPs win out when not compute bound.
Parametric networks excel when compute bound.
Softmin blending is less divergent than Argmin lookups.
A new double-covering parametrization
Expands representation capacity
Probabilistic blending
Improves gradient flow and scaling
New loss functions
Increases fidelity, stability and recall
Faster and improved multi-view stabilization
Reduce training times while improving multi-view stability
Coarse-to-fine training schedule
Improves convergence and multi-view generalization