Research
What makes the painting feel 3D — depth and light
Question: which technology makes digital oil paint read as a raised, physical surface, and not a picture of one?
The short answer
Depth is about 20% geometry and 80% light. A perfect heightfield under a flat, fixed light looks like embossed plastic. A modest heightfield with the right surface response and a light that moves with the viewer looks like paint. So the stack below is ordered by impact per unit of cost, and most of the value lives in layers 2, 3 and 6.
There are also two different kinds of depth, and oil painting has both:
- Relief depth — Rembrandt. The paint is physically raised and catches real light.
- Optical depth — van Eyck. Light passes through thin coloured glazes, bounces off a light ground, and comes back out. The colour glows from inside.
Most impasto apps only attempt the first. Both are required; the second one is cheap once mixing is done in Kubelka–Munk (see Mixing).
The stack
Layer 1 — Heightfield → normals (foundation, cheap)
One height channel per texel, updated by the paint simulation. Normals are derived per frame
from finite differences (Sobel) — never stored. Everything else sits on this.
Layer 2 — An oil-specific surface model (highest impact per cost)
Oil paint is a pigment body under a clear oil film. Render it as two layers:
- a diffuse base — the colour, whose albedo comes from the Kubelka–Munk pigment mix
- a dielectric clear coat — the linseed binder, which carries the specular highlight
The crucial detail: the coat's roughness varies per texel with binder content and age. Fresh,
fat paint is glossy. Lean or dry passages go matte — painters call it sinking in. A thick
impasto peak is glossier than a thin scumble dragged across tooth. Varying gloss across the
surface is what makes something read as oil rather than plaster, and it's almost free: the
simulation already tracks a binder / wetness channel.
Layer 3 — Anisotropic highlight along the stroke (the Van Gogh cue — our hypothesis)
Bristles leave parallel grooves running along the direction of the stroke. A grooved surface reflects anisotropically: the highlight stretches across the grooves, the way it does on brushed metal. So:
- store a stroke tangent per texel at deposit time (the direction the brush was moving)
- use it as the tangent for an anisotropic GGX specular lobe
This is the cue that makes a stroke read as made by a hand moving in a direction. We found no digital painting app that does it. It is a hypothesis — validate it against the raking-light photographs of real strokes from Track B1 before building on it.
Layer 4 — Self-shadowing and parallax (depth when the view changes)
Ridges should cast shadows into valleys, and thick paint should shift as the view angle changes.
- Parallax occlusion mapping (POM) with self-shadowing runs on every Metal device. This is the real-time answer.
- Horizon-based ambient occlusion baked from the heightfield per tile, recomputed only when a tile changes. It darkens valleys so thick paint looks heavy even under flat light.
- True displaced geometry via mesh shaders only exists on Apple GPU family 9 — A17 Pro and later iPhones, M3 and later. Use it for hero renders, the AR hang and exported video. Don't put it in the live painting loop.
Layer 5 — Glazes that glow (optical depth, van Eyck)
Kubelka–Munk has a closed-form reflectance for a translucent layer of thickness d over a substrate. Treat every dried layer as substrate for the one above it, and a thin glaze over a light passage glows the way it does on a real panel. The same model makes a glaze over a dark passage almost disappear. It costs one extra evaluation per texel, and only where layers exist.
Layer 6 — Where the light comes from (this is where "tangible" lives)
In rising order of ambition:
- Device attitude → light direction.
CMMotionManagerdevice-motion attitude (quaternion), low-pass filtered. Model it as the lamp stays fixed in the room and the painting moves, not as a light that follows the phone. Tilt it and the ridges catch. - Head tracking → view direction. A specular highlight depends on where the viewer is, not just the light. ARKit face tracking (the TrueDepth camera) gives the head position, so the highlight slides across the paint the way it does on a real object as you move your head. Without this, tilting moves the light but the eye stays fixed, and experienced viewers feel the difference.
- The real room's light. ARKit's face-tracking session also provides a directional light estimate: primary light direction, intensity, and spherical-harmonics coefficients. Brightness and colour temperature update 60 times a second. The front camera sees the room from the painting's point of view, so this is the light that would actually fall on a canvas held there. One ARKit session gives both the view vector (2) and the room light (3). It's opt-in: it needs camera permission, costs battery, and "why does a painting app want my camera?" has to be answered in the UI.
- Default studio light: a small prefiltered HDR environment for image-based lighting, used when nothing else is on.
Layer 7 — HDR gloss
Render the specular lobe into extended dynamic range (CAMetalLayer with
wantsExtendedDynamicRangeContent, rgba16Float, extended linear P3). On an XDR iPhone or iPad
display, a wet highlight is then physically brighter than paper white, which no SDR app can match.
That's what makes paint look wet. Use it sparingly — only peaks, never the base colour.
Layer 8 — Object cues (cheap, cumulative)
The canvas edge has thickness, wraps around a stretcher, and throws a soft shadow onto an implied wall when tilted. Each is small, and together they make it read as an object.
The output is a relightable object — and the museum world already agrees
- RTI (Reflectance Transformation Imaging), from HP Labs' Polynomial Texture Maps (Malzbender, 2001), is how conservators capture paintings: many photos under different light angles, then an interactive viewer that relights the surface. They use it specifically because impasto disappears in a flat photograph. Our export should be a native equivalent: albedo + normal + height + gloss maps, shown in a small WebGL viewer that reads the recipient's phone orientation.
- Relievo (Van Gogh Museum + Fujifilm, "Reliefography") 3D-scans every ridge of a Van Gogh and prints a full-scale relief replica. It's precedent that the relief is part of the artwork, and that people pay for it physically. It's also the model for a relief-print export.
What to test (feeds spike 1)
| Test | Pass if |
|---|---|
| Layers 1+2 only vs 1+2+3 (anisotropy), shown blind to painters | They pick the anisotropic version as "more like paint" |
| Gyro light alone vs gyro + head-tracked view | The head-tracked version is preferred, and worth the permission prompt |
| Room-light estimate on vs off, in three real rooms | The painting visibly "belongs" to the room (e.g. warm lamp → warm highlights) |
| POM + AO frame cost on the oldest target iPhone | Under ~2 ms at 1–2k² |
Sources
- RTI — Cultural Heritage Imaging · RTI in conservation — Conservation Wiki · NGA: RTI vs photogrammetry for painting surfaces · Smithsonian MCI on RTI
- Relievo — Fujifilm · Hyperallergic on the Van Gogh Museum 3D prints
- Rembrandt's impasto — ESRF · Natural Pigments on Rembrandt's impasto
- ARKit light estimation — Andy Jazz · AppCoda ARKit light estimation
- Apple: GPU advancements in M3 and A17 Pro (mesh shading, ray tracing) · AppleInsider explainer
- BRDF and gloss measurements · BRDF models for glossy surfaces — ACM TOG