Back to case studies

Published on August 15, 2026

Same Tree, No Shared Pixels: Holding Individual Identity Through a Season

Same Tree, No Shared Pixels: Holding Individual Identity Through a Season

The Brief

A major Japanese general contractor maintains tens of thousands of trees: along its own sites, in the parks it builds and hands over, on the streets its works touch. A certified arborist (樹木医) inspects them for bark condition, vigour, fungal fruiting bodies and form, and the register that judgment feeds is where liability sits when a limb comes down.

An inspection programme earns its keep only if this year's entry lands on the same row as last year's. Every claim an arborist makes is a claim about one individual, and a claim that cannot be attached to an individual is a photograph with a date on it. So the question underneath the whole programme is the plainest one available: which tree is this? We answer it as a localization problem. Recover where the camera stood and where it was pointed, inside the site's own metric frame, and the tree in view follows from the geometry.

We built the answer as the identity layer of YODO Asset World, and we built it on the hardest class in the family. Trees grow, they are cut back, they stand in rows and look like each other, and between August and February their appearance is replaced entirely while their identity has to hold. The site is a park in central Tokyo, surveyed once with a backpack LiDAR: 247 million points, 10,100 posed photographs at 6608×8811, two sessions registered into one metric frame with a median residual of 6.3 cm, and 564 trees registered in it as individuals, each with a permanent ID, a measured position and a dated history.

The published twin: the park's point cloud in its own colour, with ARS-T037 marked at its measured height

The sections below follow the chain YODO Asset World is built from, in order, on this one site. Localization binds every object to a row of the register by where it stands. Detection recognises the same individual on the next visit, with another device at another distance and angle, so that captures accrue into a record. Forecasting reads condition off the rounds a site already walks and returns dated work. Grounding reads each number off the 3D geometry, so the number and its evidence arrive together.

Appearance Changes With the Season

Nearly every system that recognises a specific object recognises it by how it looks. That works when appearance is a property of the object. On a deciduous tree it is a property of the month.

We can measure what that costs, because the failure has a precise location. Localizing an arbitrary photograph against the site's map is a two-step process: retrieval proposes candidate places, then a matcher verifies one. Our own capture frames retrieve at a top-1 similarity of 0.40 to 0.52. A photograph from any other camera retrieves at 0.08 to 0.36. For those photographs the right place tends to be missing from the candidate list, so the matcher has nothing to verify, and retrieval consensus clustering, which we implemented and measured, can only filter the list it is given. Foliage makes it worse in three ways at once. It is self-similar, it moves between capture and query, and the same park in August and February can share almost no pixels.

Retrieval similarity for our own frames against third-party photographs, and the published REMIND ablation

The published literature reaches the same wall from the other side. REMIND's ablation is the cleanest statement of it we have found: remove the context channel so that identity rests on the object's appearance, and IDF1 falls from 90.35 to 55.56 on the paper's own benchmark and from 62.47 to 37.53 on ScanNet++, with an ID switch rate of 51.73 %. Among neighbours of the same class, appearance alone swaps identities.

Why a Position Cannot Name a Tree

The obvious repair is to bind each observation to a coordinate. A phone's GNSS is good to about 3.5 m in the open and worse under a canopy. Whether that is enough is a question about the site.

So we asked the site. Across the 564 registered individuals the median distance to the nearest other individual is 3.10 m. Six in ten of them (59.9 %) have another registered individual inside a 3.5 m radius. Inside 10 m the average tree has 6.3 neighbours, and the most crowded fix in the park covers four more trees besides the one it was meant to name.

The 564 registered individuals in plan, the nearest-neighbour distribution against a 3.5 m fix, and what a fix leaves out

A survey-grade receiver puts the operator inside a centimetre and reports where the operator stood. Which trunk is in the frame depends on where the lens was pointed, and a position fix carries no orientation. A camera pose has six numbers, and a fix supplies three of them.

The missing three decide the answer. At 7 m one degree of aim moves the ray 12 cm, and a heading error of 24° lands it on the neighbour 3.10 m away. Staying nearer the intended trunk than the next takes heading good to about 12°. At 20 m, where a street tree is photographed from the pavement, the same gap leaves 4.4°.

Localizing the Photograph

So we localize the camera, and let the trees follow from it.

A photograph is localized inside the twin, which returns its 6-DoF pose in the same metric frame the assets are registered in. Position and orientation arrive together, out of the image, because both are what the image constrains. Once the camera is placed, the question of which assets the photograph contains is settled by projection and a visibility test: frustum, range, occlusion. Correspondence becomes analytic where recognition is probabilistic, and it does not degrade when two objects look alike, because looking alike was never part of the calculation. The ID follows from where the camera was standing and where it was aimed.

What the method needs is the photograph. No GNSS, no RTK, no control points, no markers on the trunks, and no particular camera. The clip below was shot on an ordinary phone, in one hand, walking.

Three frames of the association layer's overlay on a smartphone clip, with registry IDs on the trunks

The numbers as they stand today. On 40 held-out frames from our own rig the pose lands at a median 1.2 cm and 0.16°, an order of magnitude inside the 0.2 m and 2° that bounding-box projection needs at 2 to 10 m. On a smartphone clip walked through the park, 18 of 31 keyframes localize. Through those poses the association layer assigned 7 tracklets to real asset IDs with zero wrong IDs, re-pinning two fragments of one tree to the same ID, leaving 6 undecided and abstaining on the deep-canopy tail.

In that run, correcting a single number moved assignments from 4 to 7 and the top tracklet's evidence from 3.6 to 32.6. The number was the query focal, assumed at 1300 px repo-wide while the clips were shot on the phone's 0.5× ultra-wide at 806 px.

What bounds the 18 is the problem the first section described: a photograph from another camera fails at retrieval. The designed fix works on the reference side. Part of the matching moves onto 3D structure, which a season leaves alone, and the reference set is densified with viewpoints nobody walked, rendered from a 3DGS model in the same metric frame. That model exists and is already in the viewer: 7.97 million gaussians, aligned to the LiDAR at zero translation offset.

One held-out view: the photograph, the stock trainer collapsed to fog, and the best model so far with the near bench missing

Two of its failures were invisible in the literature, and we root-caused both. The stock trainer preset walks weakly observed gaussians below the cull threshold until the scene collapses to fog; turning one flag off is worth 8.6 dB. Separately, the SLAM poses disagree between adjacent views by 0.195°. Bundle adjustment cut that sevenfold and moved the fine-detail metric 4.2 times, after a dozen model-side levers had moved it by less than 0.01. A second season of capture adds the other half of the year to the reference set.

The Gates That Decide Whether We Answer

A wrong pose poisons everything downstream, so the system is built to stop. The numbers that decide whether it speaks live in one file, each with the measurement that set it, and they are calibrated on the easy set, held-out frames from our own rig, and frozen before the honest set, photographs from other cameras, is touched. Re-tuning a gate after seeing held-out data turns a held-out measurement into a fitted one.

The five frozen decision gates, and route coverage against query resolution

Three findings came from testing what the gates miss.

A gate that passes does not validate what it did not measure. Above 60 inliers the pose error sits at a 4 cm median; below 30 the P90 opens to 21 m, which is why 30 is the abstain line. It says nothing about the focal: 37.5 % of deliberately wrong-focal pairs still clear it, and one frame returns 53 inliers at 2.9 times the correct focal. So the focal is swept, never assumed.

A second criterion has to be independent of the solver. We documented reprojection RMS as a quality gate, then measured it and found it inverted. The single wrong pose in the easy set has an RMS of 0.00 while the 39 correct ones sit between 2.53 and 4.72, because RANSAC selects the inliers and then the residual is measured over those same inliers. Few inliers, near-exact fit, low RMS. Low RMS flags over-fitting. What works is drift against an interpolated pose, which compares against information the solver never saw.

More pixels is worse. Feeding the same 31 keyframes at five sizes in one batch, 720 px localizes 19 and 1600 px localizes 14. In a canopy-heavy scene the extra resolution buys leaf detail, the keypoint budget spends itself on foliage, and the trunks, kerbs and buildings that carry stable structure go unsampled. Raising the budget recovers one frame.

The same discipline retired a change we wanted. Reading the rig's second camera head lifts route coverage from 18 to 27, and worsens easy-set median translation error by 39.8 % across 26 of the 36 shared frames. Its own pre-registered rule said no.

The gates show up where a person can see them. For every clip it is handed the product reports two counts: how many assets the frames were seen to contain, and how many it was willing to name. The gap between them is the abstention, and nothing is written until a person confirms it.

The product surface: a clip is handed in, IDs arrive with a seen and matched counter, and a person decides whether to write it to the ledger

The same counter runs live in camera mode: point a phone and the ID lands on the trunk as it records. There the constraint is stability. A label that flickers between two trees is worse than one that takes three seconds to appear, and much worse than an honest question mark.

Four Keys, and the Two Ways They Break

Everything above assumes an asset stays where it was registered and only the background changes. Dropping it changes the kind of failure. A background that changes costs coverage, and the system says so by abstaining. A target that has moved costs correctness: nearest-position matching returns a confident wrong ID when another object drifts into the old coordinate, no existing gate fires, and a data model that stores position as a field has no way to notice that a coordinate has gone stale.

Two assumptions break, and each gets a replacement. Position becomes an observation stream, with the 2026 survey backfilled as the founding epoch-0 event so that nothing already measured is thrown away. Identity becomes a two-stage assignment over four keys: a position prior, an appearance embedding, a local-feature re-rank, and neighbour context.

The four identity keys against seasonal change and relocation, and the two-stage assignment

Two stages, because of an asymmetry that a single weighted sum hides. A relocated asset loses its position key and its neighbour key in the same instant, since both describe where it used to be. Sum all four and the cost of "this thing moved" becomes indistinguishable from the cost of "something else is standing here now". So the first stage assigns everything that did not move, using all four keys over each co-visible neighbourhood, and emits as a by-product the evidence that something expected is absent. The second stage takes the residue, lets appearance and local features lead, enumerates four hypotheses (relocated, new, departed, ambiguous swap) and sends anything above the cost threshold to a human queue.

Two same-model objects exchanged between visits are indistinguishable to any visual system, ours included, so such assets carry model-level identity and the record says so.

What a Patrol Can Write Back

Identity is the precondition for the thing an inspection programme actually wants. The national guideline for inspecting and diagnosing urban park trees was revised on 2026-03-30 to recommend folding routine inspection into ordinary patrols, and digitising records to lighten the work of keeping the register current. The inspection tools in use today, the diagnosis model this programme already runs among them, begin with a photograph of a tree someone has already identified. With identity resolved automatically, a keeper's normal round, filmed on a phone and annotated with at most a spoken sentence, becomes the input.

Update is several jobs. Treated as one, it produces a number that looks updated and is old. Passive capture can write evidence (when, from what angle and at what quality an individual was seen, which on its own is a record of inspection), condition and findings, and disappearance. Precise measurements such as girth and height stay as surveyed. Geometry sits between the two. A deletion is free, a moved object keeps its original LiDAR points and only its pose is estimated, and genuinely new geometry needs a deliberate orbit video. A patch that misses the cloud's own accuracy stays out.

What an ordinary round can update, the evidence event written back, and the three-state receipt

What gets written is an evidence event, (asset_id, time, claim, evidence, confidence, provenance), appended to the register and committed by a person. A keeper's remark that a tree is leaning is the highest signal-to-noise input the system receives, and it can be stored only because there is an ID to attach it to.

The hard half is proving that nothing changed. Nine tenths of inspection labour goes into confirming that there is no anomaly, and passive observation cannot by itself separate "I did not look at it" from "it did not change". The change-detection papers we surveyed report precision and recall on change and give no confidence for the unchanged. So a round is designed to return a receipt per asset in three states: changed, confirmed unchanged, not seen clearly. Confirmed unchanged is a calibrated probability resting on a criterion independent of the fitted residual, the same rule the localization gates follow. Only structural change triggers anything (removed, added, moved); season and light trigger nothing, and a single photograph can only flag an asset.

The two LiDAR sessions, captured nineteen minutes apart, give a zero-change control set for measuring the false-alarm rate, and positive cases are synthesised by deleting or moving a registered tree inside the model.

Forecasting From the Rounds Already Walked

The same rounds are where forecasting starts. Condition forecasts come off the patrols a site already makes, with nothing instrumented: project the register forward in time and it becomes a prediction. The expected state of an asset today is a robot's change-detection prior; two years out it is a pruning plan; conditioned on a typhoon's wind field it is a risk score.

The starting input is already in the register: a screening tier for 355 individuals, read from their photographs, with 266 healthy, 60 needing attention, 5 suspected dead and 24 not assessable from the views on file. An inspection-priority ranking and a typhoon watchlist build on it, and what comes back from either is a dated work item: which individual, which check, by when.

The forecast stays inside what the rounds have observed. Its outputs are screening, every value carries a calibrated uncertainty or abstains, and the post-storm rescan that grades a watchlist is also the label set that trains its successor. Each round adds to the record the next forecast reads from.

Where Measurement Comes From, and Where Class Does

Every number in the register is read off the geometry itself. Girth comes from a band of points at a recorded height above the fitted ground, height from the column above it, crown and lean from the same individual's points, and each value keeps its capture, date and fit residual. The height audit draws the measured line back onto the tree's own photograph, so the number and its evidence arrive together. The chat answers from these values: a question in Japanese, English or Chinese calls eleven tools over SQL, every number in an answer is read from SQL, and 52 frozen evaluations check that it stays that way.

One individual's record, a measurement with its error bar derived on the spot, and a question answered from the register

Registering 247 million points as individuals raised a question we expected to be easy. Which of these columns is a tree?

Class is not recoverable from this point cloud. Eleven stored attributes, seven column and crown features and fifteen radiometric features all overlap between objects as different as a tree and a lamp post, and the scanner has no multi-return, so the standard foliage-versus-solid discriminator is unavailable. The rule that came out of it is now project-wide: the photograph is the authority for class, the point cloud for measurement.

The split of authority, and what it took before a model's verdict could change a row

Which moved the difficulty into a second question: how do you make a vision-language model's verdict safe enough to change a row in a register? We designed it the obvious way first, with one model's verdict flipping an asset's status automatically. We pre-registered a bar of 5 % for the order-effect flip rate, measured 35 %, found the verdicts uncorrelated with the point-cloud reference, and did not ship it.

What shipped requires independent judges, unanimity among everyone who looked, and retirement on the photograph alone. On 2026-08-13 it retired 62 assets whose photographs contradicted their registered class, taking the active count from 626 to 564, and restored two that an earlier and weaker protocol had retired when a second judge saw a real trunk on the sheet. Five hidden anchors with known answers rode in every batch: on trees the judges went 12 for 12, and on the three negative anchors no judge ever said tree while none would confirm the negative either. They abstained, at 52 % for the subagent judges and 20 % in session.

Nothing was deleted. IDs, photographs, measurements and history survive retirement, every row carries a changelog entry naming its evidence, and the registry rebuilds from scratch to the same active set.

What Accrues

Perception gets cheaper every month and arrives free. A ledger of the same individuals across years arrives only by being kept, one visit at a time, which is what makes it worth starting early and worth owning afterwards.

Every epoch also compounds into something with no public equivalent: a supervised corpus for instance re-identification of outdoor objects whose appearance is replaced by time. Every instance-persistence benchmark is indoor, the outdoor change datasets carry no identity, multi-temporal forestry data has no casual capture, and the combination of outdoors, cross-season, cross-device and instance-level is unclaimed. We are among the few holding data that can build it, because the capture is ours.

And the layer underneath is class-blind by construction. Nothing in the ID, the evidence, the localization or the association knows what an object is, because correspondence is solved by pose and projection. Two things are per class: the geometry that cuts individuals out of the survey, and the text prompt the 2D detector is given. A bench, a lamp post, a guardrail, a valve or a sign enters the same ledger on the same terms as a cherry.

It is sensor-blind on the same principle. What the registry asks of a survey is that the result be metric. A backpack LiDAR built this twin; a handheld 360 camera walked through a site yields the same shape of product, geometry with a measured scale that individuals can be cut out of, and we have run that pipeline end to end on another site.

The primary reader of this ledger will be a machine. A robot on site needs three answers. Where am I: SLAM supplies that, relative to itself. What is around me: perception supplies that. Was this here yesterday, and is this change expected: only a registered prior can answer, and the inspection products we surveyed keep their change baselines in their own clouds under their own IDs, while localization services return a pose with no identity, history or rules attached. The ledger is read-write, so what a robot captures flows back in as evidence, and working the site becomes collecting it. That is the route: asset intelligence for people now, then the same ledger as the context layer that people, agents and robots all work from, and, once every job on site reads and writes through it, the operating system for physical work. The register answers people through the chat today.

What We Took Away

Recognising a tree is now a commodity. What an inspection programme needs on top of it is to know that this tree is the one it looked at fourteen months ago, and to say so with a stated confidence or abstain. That takes a register kept over time, and each round adds to it.

The arborist's judgment is what the contractor pays for. Identity turns a sequence of those judgments into a record that carries forward from one year to the next.