A Radiance Field That Knows Which Blob Is the Apple
Semantic Radiance Fields (SRF) take a pile of posed photos, ask SAM 3 what is an apple / branch / leaf, and bake those labels into a 3D radiance field. The paper (arXiv:2608.13095, 13 August 2026) is an oral at the IJCAI 2026 Spatio-Temporal Reasoning and Learning workshop. Authors are at Leipzig University, the Systems Research Institute of the Polish Academy of Sciences, and Wrocław University of Economics.
The point is not a prettier NeRF. It is a simulator you can query: render a new camera, ask “is this point an apple?”, ask “is this point solid?”
What they actually built
They extend FruitNeRF from one semantic channel to C independent binary heads. A point can be apple and leaf. Semantics do not back-propagate into geometry, so the tree does not collapse onto class edges.
The example scene is FruitNeRF’s apple tree: 311 posed frames at 6000×4000, downscaled 4× for training. SAM 3 is prompted separately with “apple,” “branch,” and “leaf.” Training: 500,000 iterations, batch 4,096 rays, Adam, about 4 hours on one NVIDIA H100.
A trained field exposes three calls:
- Render(pose) — RGB, semantic map, depth
- Semantic(x) — per-class probabilities at a 3D point
- Occupancy(x) — density for collisions
The apple-reaching sketch
They outline (they do not run a full RL study) an orchard reaching task. MuJoCo would own rigid-body dynamics. The SRF would render the wrist camera and supply occupancy. Reward: get the gripper near a fruit. Collision with the branch class ends the episode.
The authors say the same lifting could move to 3D Gaussian Splatting for faster rollouts, and that a time axis would make the field spatio-temporal.
A Human’s Take
Training a picker in a fake orchard is easy. Training it in a field you scanned last Tuesday is the trick. I like that they keep SAM 3 honest by not letting semantics rewrite the geometry. Now run the policy. A four-hour H100 bake is fine if the arm actually finds the fruit.