← Back to Deep Dives

When Photos Become Places

NeRFs and Gaussian splats can turn ordinary photos into scenes we can move through. The results feel like 3D photography, but what is being captured is stranger—and more useful—than a conventional model.

A warm editorial scene showing overlapping photographs transforming into a navigable three-dimensional room made of glowing points and translucent Gaussian splats.

A photograph usually tells you where to stand. The photographer chose the spot, pointed the camera, and kept one flat view of what was there. Video adds time, but it still chooses the path for you.

\n

Over the last few years, that bargain has started to change. With enough overlapping photos or a carefully recorded phone video, software can now produce something that feels less like a picture and more like a place. You can move the camera after the capture is finished. You can look around an object, drift through a room, or revisit a location from an angle that was never recorded as a normal photograph.

\n

The results can be startling. A rough phone capture can preserve soft shadows, foliage, and little imperfections that would take a skilled 3D artist a long time to rebuild by hand. Reflections may look convincing too, though they often reveal the illusion when the viewpoint moves too far. It can feel as if the scene was somehow lifted out of the real world.

\n

That is where terms like photogrammetry, NeRF, and Gaussian splatting start appearing. They often get grouped together because they can produce similar fly-through demonstrations, but they are not three names for the same thing. They represent different answers to a much older question: how can a computer use a collection of images to understand enough about a scene to show it from somewhere new?

\n\n

The shared starting point is easier to see than the terminology suggests. Each method begins with overlapping views and an estimate of where the cameras were. The important difference is what happens next: photogrammetry tries to build surfaces, a NeRF learns a field that can answer viewing questions, and Gaussian splatting fits explicit translucent primitives that can be drawn quickly.

\n\n
Process diagram showing overlapping source photos becoming camera positions and a sparse point cloud, then branching into a photogrammetry mesh, a NeRF sampled by camera rays, and explicit 3D Gaussian ellipsoids before producing a rendered novel view.
The methods share a capture and camera-estimation stage, then diverge in what they build and how they produce a new view. Select the diagram to enlarge it.
\n\n

Before the Splats

\n

The idea of building 3D scenes from photographs did not begin with modern AI. Photogrammetry has long been used in surveying, mapping, archaeology, and visual effects. In a common workflow called Structure from Motion, software looks for the same recognizable points across multiple images. If a corner, crack, window, or patch of texture appears in several photographs, the software can compare how its position changes, estimate where the cameras were, and place that point in 3D space.

\n

Repeat that process enough times and a sparse cloud of points begins to form. From there, another stage can build denser geometry and eventually a textured mesh: the kind of object most people imagine when they hear “3D model.”

\n

This works because moving a camera reveals depth. Hold one finger in front of your face and alternate between closing your left and right eye. Your finger seems to jump against the background. A set of photographs taken from different positions gives reconstruction software a more complicated version of that same clue.

\n

The strength of photogrammetry is that it tries to recover actual surfaces. That makes the result useful in conventional 3D software, where a mesh can be measured, edited, simplified, given new materials, or used for collision in a game engine.

\n

The weakness is that the real world does not always cooperate. Blank walls offer very few distinctive points to match. Glass and water show different things from different angles. Reflections move. Leaves and people move. Thin objects disappear. The underside of a chair cannot be reconstructed if the camera never saw it. Photogrammetry can be excellent, but it asks the images to provide convincing geometric evidence.

\n\n

NeRFs Change the Question

\n

When the original Neural Radiance Fields paper appeared in 2020, the most interesting change was not simply that the images looked better. NeRF approached the scene differently.

\n

Instead of immediately trying to produce a clean surface for every object, a NeRF learns a function that can answer a visual question: if a camera looks through this point in space from this direction, what color and density should it encounter?

\n

That sounds abstract, but the practical goal is familiar. The system studies a set of photographs and adjusts itself until it can reproduce the views it was shown. Once it has learned the scene well enough, it can render views from camera positions that were not part of the original set.

\n

The result is sometimes described as a volume or field rather than a normal model. There may be useful geometry hiding inside it, but the main job is to recreate how the scene appears.

\n

That distinction matters. A NeRF can produce a beautiful new view without giving an artist the tidy surfaces, separate objects, editable materials, and clean topology expected from a traditional asset. It is a little like learning to convincingly describe what a room looks like from anywhere inside it without first producing a perfect architectural drawing of the room.

\n

The original results were impressive enough to make NeRFs one of the defining computer-vision developments of the early 2020s. They were also slow. Training a scene could take hours or longer, and rendering new views involved asking the learned field many questions along every camera ray.

\n

Research moved quickly. In 2022, NVIDIA's Instant Neural Graphics Primitives, better known as Instant-NGP, showed that a different way of encoding spatial information could cut training dramatically. NVIDIA reported real-world NeRFs trained in under five minutes on its test hardware, with some neural-graphics tasks converging in seconds. NeRFs did not suddenly become effortless, but they became much easier to imagine as practical tools rather than only research demonstrations.

\n\n

Why Gaussian Splats Felt So Immediate

\n

Then, in 2023, the Inria-led paper “3D Gaussian Splatting for Real-Time Radiance Field Rendering” rearranged the tradeoffs again. Its authors reported high-quality novel views at real-time rates on their test scenes and hardware. The name is intimidating, but the visual idea is friendlier.

\n

Imagine a scene represented by millions of tiny, translucent blobs floating in 3D space. Each blob has a position, color, opacity, size, and orientation. They are not all round. Many are stretched or flattened into small ellipsoids so they can settle along surfaces and fill the view efficiently.

\n

When the scene is displayed, those Gaussians are projected onto the screen and blended together. Seen individually, they can look like a cloud of colorful fuzz. Seen from the views the scene was trained to reproduce, they combine into a remarkably convincing image.

\n

The original 3D Gaussian Splatting method still begins with familiar reconstruction work. It commonly uses camera positions and a sparse point cloud estimated by software such as COLMAP. It then creates and adjusts the Gaussians, adding detail where the scene needs it and removing primitives that are not helping.

\n

What made the result feel different was speed. NeRFs had already shown that view synthesis could preserve details traditional reconstruction sometimes struggled with. Gaussian splatting made similarly photorealistic scenes much easier to render interactively because the result could be drawn through a fast rasterization process rather than repeatedly querying a neural network along every ray.

\n

This is why the “NeRFs versus splats” story became so tempting. One approach was associated with slow neural rendering; the other arrived with real-time demonstrations that worked in a browser or game-like viewer. But that version is too neat. Gaussian splatting did not eliminate the camera-calibration work inherited from photogrammetry. It did not turn every capture into a clean mesh. It did not make NeRF research stop. It created a particularly useful point in the design space: high visual quality, fast rendering, and a representation made of explicit primitives that can be inspected and manipulated more directly.

\n\n

What My Desk Capture Actually Showed

\n

I saw that gap in a small desk experiment. I captured overlapping views of a U-shaped desk and the room around it, reconstructed a Gaussian-splat scene, and then rendered the 20-second path below. This was a practical observation, not a benchmark: one room, one capture, one reconstruction, and no controlled comparison against a NeRF or photogrammetry pipeline.

\n

Along camera positions close to the source views, the desk looks coherent enough to feel solid. As the camera keeps moving into poorly observed space, the scene opens into translucent fragments. Thin edges smear. Gaussians stretch across gaps. Parts of the room that the camera never saw do not magically become complete.

\n

That failure is the point. A captured scene can look nearly photographic from a familiar path and still fall apart when you move too far outside it. Reflections may also appear tied to the captured views instead of behaving like surfaces under newly simulated light.

\n
The captured desk looks coherent along familiar viewpoints, then opens into translucent fragments as the camera continues beyond the best-observed area. Select play to load the 20-second video.
\n

People sometimes call the stray fragments floaters, which is an unusually accurate technical term. They are pieces of the reconstruction that look plausible from one angle and obviously wrong from another.

\n

Memory is another tradeoff. Millions of explicit Gaussians can be fast to draw but expensive to store and move. That matters when the goal changes from viewing one captured room on a workstation to streaming a neighborhood to a phone or a pair of glasses.

\n

Editing is improving, but it is still useful to ask what kind of editing we mean. Removing a cluster of splats is not the same as selecting a cleanly modeled chair. Changing captured colors is not the same as assigning a physically accurate material that responds correctly under new lights. Extracting a mesh is possible in newer workflows, but it does not guarantee the sort of topology an artist would have built deliberately.

\n

None of this makes the capture a trick. It means that “looks real,” “contains accurate geometry,” and “works like a conventional 3D asset” are three different goals.

\n\n

The comparison below makes that gap visible. The intermediate representation can look like a mesh, a learned volume, or a cloud of ellipsoids, while the finished render smooths those differences into a convincing image. The failure views are just as useful as the polished one because they reveal what the capture never truly recovered.

\n\n
Visual comparison of reality-capture stages and real outputs, showing source imagery and camera poses, photogrammetry geometry, a NeRF volume, Gaussian splats viewed as colored ellipsoids, a polished rendered scene, and examples of artifacts outside well-captured views.
A convincing rendered view can hide very different underlying structures—and weaknesses that become obvious when the camera leaves well-observed territory. Select the comparison to enlarge it.
\n\n

The Field Did Not Pick One Winner

\n

Technology stories often become simpler after the fact. A new method arrives, replaces the old method, and the timeline moves forward. Reality capture is developing in a messier and more interesting way.

\n

NeRF researchers kept working on faster rendering, larger spaces, and better editing. Google's SMERF project explored large scenes that could render in real time on everyday devices. Nuvo explored ways to give neural fields editable 2D texture maps instead of leaving appearance locked inside an opaque representation.

\n

Gaussian research spread in several directions at once. Some projects try to recover cleaner surfaces. Others add semantic understanding so parts of a scene can be found and selected by meaning. Some separate lighting from material appearance so the capture can be relit. Compression work tries to make splats practical to store and stream.

\n

Time is becoming part of the representation too. A static 3D capture lets the camera move through a frozen moment. Research such as the CVPR 2024 project 4D Gaussian Splatting adds change over time, allowing dynamic scenes to be rendered from different positions as they unfold.

\n

That phrase describes a family of developing techniques rather than one settled standard, but the direction is easy to understand. A photograph saves a view. A video saves a sequence of views chosen by the camera operator. A dynamic volumetric capture tries to save enough of the event that the viewer can choose the viewpoint later.

\n

That is a much larger file and a much harder problem. It is also where the line between video and 3D begins to feel less stable.

\n\n

From Demonstrations to Infrastructure

\n

The next stage may be less visually dramatic than the first splat demos, but it could matter more. On February 3, 2026, the Khronos Group announced a release candidate for KHR_gaussian_splatting, an extension that stores Gaussian splats in glTF. A release candidate is open for feedback before ratification, so this was not yet a final standard. The goal is still important: give compatible tools a shared baseline for describing splats instead of forcing every application to invent its own version.

\n

Open representation and compression formats will not solve capture quality, missing geometry, or lighting. They can make the results easier to move between viewers, editors, mapping systems, and devices. That is how a striking research method starts becoming part of ordinary infrastructure.

\n

The practical uses span several lanes, though their maturity varies. The Khronos announcement points to current work in geospatial capture, digital twins, infrastructure, and large-scale visualization. Visual-effects teams can preserve the appearance of real locations. Museums and preservation groups can record places and objects that are difficult to move. Researchers are also exploring richer visual representations for robotics, games, and spatial interfaces.

\n

Those uses do not all need the same kind of scene. A film shot may prioritize appearance. A robot cares about stable position and navigable structure. An architect may need accurate measurements. A game designer may need clean collision and objects that can be edited independently.

\n

The question is not only which method looks best. It is which representation preserves the information the job actually needs.

\n\n

What It Feels Like We Are Saving

\n

There is still something strange about seeing a Gaussian splat for the first time. The scene looks solid until the camera moves beyond the safe area and the illusion opens up into points, haze, and stretched fragments. It is both more and less than a model. It can preserve details a handmade asset would miss, while lacking surfaces a normal model would take for granted.

\n

That tension is part of what makes the technology interesting. We are getting better at capturing how a place appears before we have fully solved how to turn that appearance into a clean, editable, relightable world.

\n

NeRFs were an important step because they showed how much of a scene could be recovered by learning to reproduce its views. Gaussian splatting mattered because it made that kind of captured scene feel fast, visible, and manipulable. The work happening now is about filling in everything the first demonstrations left out.

\n

It is tempting to call these captures 3D photographs. That is probably close enough to explain the feeling, as long as we remember how much is hiding inside the phrase. They are photographs with camera positions, reconstruction, optimization, and rendering wrapped around them. They are not perfect copies of reality, but they are no longer just flat records of where a camera once stood. They let the viewpoint move after the moment is gone.

\n\n

Bilawal Sidhu's visual explainers on reality capture helped inspire this review. The technical claims and chronology above were checked against the linked papers, project pages, and Khronos announcement; the desk capture is our own limited experiment.

\n\n\n
Ryan

Ryan

Architect of digital systems and thoughtful experiments