Jonghoon Ahn.← Research portfolio
01 · Representation · Controlled benchmark · 2026

NeuralScene Bench

Before optimizing Gaussian Splatting, I wanted to know what advantage the representation actually offered. This study compares 3D Gaussian Splatting and NeRF under the same Nerfstudio data and evaluation pipeline.

Mip-NeRF 360Splatfacto / gsplatNerfacto / tiny-cuda-nnPython 3.11PyTorch 2.7.1 + cu128A100 40GB
Open research branch on GitHub

Why this came first

I did not want to optimize a representation before understanding what I was optimizing.

Gaussian Splatting is attractive because it can reconstruct detailed scenes and use an explicit set of Gaussians for rendering. NeRF uses a different, implicit neural representation. Rather than assuming one was better, I set up a controlled comparison under a fixed nominal optimization budget.

The question was narrow: under the same data parser, held-out split, image scale, camera treatment, and 5K iteration count, how would the two representations differ in held-out reconstruction quality?

Protocol

Controlled inputs

Bonsai and Garden from Mip-NeRF 360. Both methods use the same COLMAP parser route, logical images path with explicit 2× downscale, interval held-out split, and evaluation interval 8.

Controlled camera treatment

Nerfacto defaults to camera optimization while Splatfacto defaults off in the pinned Nerfstudio version. I disabled camera optimization for Nerfacto so the comparison would not silently include different camera refinement behavior.

Backends

Splatfacto uses gsplat. Nerfacto uses tiny-cuda-nn compiled for A100 sm_80. The benchmark runs in an isolated Python 3.11 environment with pinned dependencies.

Metrics

Held-out PSNR, SSIM, and LPIPS from ns-eval. Evaluator throughput was measured during development but is not treated as headline evidence because repeated evaluation showed sensitivity to execution conditions.

Held-out results

SceneMethodBackendPSNR ↑SSIM ↑LPIPS ↓
BonsaiSplatfactogsplat27.9380.89580.1775
BonsaiNerfactoTCNN25.7960.80940.2497
GardenSplatfactogsplat24.0010.67930.3214
GardenNerfactoTCNN22.3360.50770.5235
Bonsai PSNR Δ+2.14 dB
Garden PSNR Δ+1.66 dB
Bonsai LPIPS-28.9%
Garden LPIPS-38.6%
NeuralScene Bench summary showing Ground Truth, Splatfacto and Nerfacto for Bonsai and Garden
Two-scene summary. Ground truth is shown beside Splatfacto and Nerfacto under the same nominal 5K-step benchmark protocol. Click to open the full-resolution figure.
Bonsai held-out reconstruction comparison across four views
Four held-out Bonsai views. The comparison shows Ground Truth, Splatfacto, and Nerfacto at the same camera views. Fine flower structure, bicycle spokes, cloth texture, and object boundaries make the quality difference easier to inspect than aggregate metrics alone.
Garden held-out reconstruction comparison across four views
Four held-out Garden views. Outdoor foliage, table geometry, pavement edges, and local texture provide a second scene for checking whether the same reconstruction-quality trend remains visible.
What I found: the reconstruction-quality trend was consistent across both tested scenes.

Splatfacto achieved higher held-out PSNR and SSIM and lower LPIPS on Bonsai and Garden under the fixed nominal 5K-step protocol. This is evidence about quality at that budget, not a claim that 3DGS universally outperforms NeRF or that both methods were equally converged.

What surprised me

The strongest result was not simply “GS wins.”

The important result is narrower: under the same nominal 5K-step budget and controlled data pipeline, Splatfacto reached higher held-out reconstruction quality on both tested scenes.

That made the next question more useful. If the representation was promising at this operating point, could I reduce the cost of creating it without giving that quality back?

What this led to next

Interpretation boundary

Matched 5K iterations are not matched compute, convergence, capacity, or wall-clock budget. Nerfacto is commonly trained for longer schedules, so these results are best read as a fixed nominal optimization-budget comparison. Repeated evaluator runs reproduced the core PSNR, SSIM, and LPIPS values while throughput varied, so evaluator FPS is treated as implementation-specific diagnostic information rather than a stable representation benchmark. The study currently covers two scenes, not a statistically representative scene distribution.