Back to case studies

Published on August 15, 2026

One Tree, Many Cameras: A Capture Protocol a Diagnosis Model Can Trust

One Tree, Many Cameras: A Capture Protocol a Diagnosis Model Can Trust

The Brief

A major Japanese general contractor maintains tens of thousands of trees: along its own sites, in the parks it builds and hands over, on the streets its works touch. A certified arborist (樹木医) inspects them for bark condition, vigour, fungal fruiting bodies and form. The judgment is expert and slow, and the register it feeds is where liability sits when a limb comes down.

A diagnosis model already existed for two species, cherry and zelkova, returning three condition categories from four part views per tree: root collar, trunk, canopy, whole tree. We were not asked to improve it. We were asked the question one step upstream, which is the question that decides everything: how should the trees be photographed?

That sounds like logistics and it is optics. A crew with a phone, a crew with a helmet camera, a surveyor riding past on a bicycle and a drone overhead differ in two respects that matter. How many pixels land on a square millimetre of bark, and whether those pixels moved while the shutter was open. Model accuracy, trees covered per day, and whether the programme is affordable at all sit downstream of those two numbers.

How Many Pixels Land on a Millimetre

A bark lesion reads at roughly one pixel per millimetre of bark. We took that as the working floor for this diagnosis task, and it converts straight into a rule about where a person may stand, because one millimetre at distance d subtends 0.0573/d degrees. A lens that resolves r pixels per degree delivers

px/mm = r × 0.0573 / d

and nothing about the model moves that line.

Sampling on bark against distance, for four optics, with the 1 px/mm floor

Once the curves are on the page, the field manual explains itself. It was written when a phone meant a 12MP wide lens at about 60 px/°. That lens clears the floor only inside 3.4 m, which is why the manual tells the crew to walk in and crouch at the root collar. Those instructions were never about trees. They were about a lens.

A current 48MP 4× telephoto resolves about 390 px/°. From a 7 m ring it puts 3.2 px/mm on bark, three times the floor, and from the same position the wide lens still has the whole tree. The rule "get closer" has lost its subject. A street tree can be sampled from the pavement it stands on, 10 to 20 m away, without anyone crossing the road.

The same arithmetic removed a 360 camera from the plan, where an earlier version had put it at the centre. At 43 px/° it clears the floor only within 1.5 m, and under a canopy it drops into a 4-in-1 mode at about 21 px/°, a limit of 1.2 m, with no raw file to recover from. The intended use made it worse. Orbiting the trunk at 1.5 m sweeps the stitching seam across the bark over and over, in the near-field geometry where stitching artefacts are documented. A seam on bark is a lesion that was never there, or one that has been wiped away.

Whether Those Pixels Moved

Walking an orbit of radius r at speed v turns the camera at v/r radians per second. At r = 2 m and a gentle 1 m/s that is 29°/s. Under a canopy the scene sits around EV 10 to 12, and at f/1.9 and ISO 800 the shutter lands between 1/60 and 1/125 s. One exposure therefore smears the image by about a third of a degree, which on a 43 px/° sensor is 12 to 15 pixels, against a target of one pixel per millimetre.

Smear during one exposure against walking speed, for two orbit radii and two shutter speeds

A third of a degree costs different lenses different amounts, and the ranking runs opposite to the previous section. Walk the 7 m ring at that same 1 m/s and the camera turns at only 8°/s, but on a 390 px/° telephoto the same exposure window gives 26 to 53 pixels of smear. At that distance on that lens it is eight to seventeen millimetres of bark drawn into a single line. The optics that solved the sampling problem make the motion problem worse, because smear is counted in pixels and a sharper lens has more of them per degree.

Two things follow.

The closer you walk, the worse it gets. Angular velocity rises as the radius shrinks, so the tight orbit that looks efficient is the one that destroys the most detail. And the place where the shutter is slowest, under the canopy, is where the trunk is.

Stopping is not a slower way to capture. It is the only way. Standing still zeroes the angular term, and it also unlocks the multi-frame stacking a modern phone performs once it decides the scene is static. Time saved by shooting on the move is paid back the moment the frames turn out unusable.

The Study

Arithmetic does not settle what a crew will actually do, so the July session was run as a controlled comparison rather than a collection. Four camera mounts, phone in hand, action camera in hand, on a helmet and on the chest, crossed with two intents and two speeds, plus a bicycle passing the tree from the front and from the side. Cherry and zelkova in a Tokyo park, including individuals an arborist had already annotated, so the model's answers could later be checked against an expert's.

The capture design matrix: four mounts, two intents, two speeds, plus bicycle passes

"Deliberate" means the operator was circling the tree for the camera. "Incidental" means they walked past while doing something else. The second one matters commercially, because it is the only condition that scales to every site visit a company already makes. A bicycle can only ever pass.

The session produced 49 videos covering 68 tree passes, and 923 frames were extracted from them against the model's part requirements. Extraction follows what the diagnosis does with the images rather than a fixed interval:

  • Root collar, trunk and canopy are scored per photograph, and a local defect such as a cavity, a fruiting body or a patch of bark loss is invisible from the wrong side. Angles are what is scarce, so one frame comes off roughly every 3 seconds of orbit, 8 to 24 per part.
  • Whole tree works differently. Vigour and form are computed from the whole-tree view alone, and when it is missing the model returns no vigour, which propagates into an absent overall rating rather than a weaker one. It gets 4 to 12 frames at wider spacing.
  • Poor frames are excluded rather than down-weighted. The overall rating aggregates across detections, so one false positive raises the whole tree's risk level. In an aggregating system, adding a bad photograph is not a neutral act.

One cherry, four bicycle passes, the four part views, with a focus measure on each tile

The bicycle rows show the blur budget in the shape it takes in the field. A pass returns the whole-tree and canopy views perfectly well, because those subjects are far away, the angular velocity against them is small, and the focus measure stays in the 7,000 to 9,000 band. The two views that carry bark are where it collapses: 953 on the fast frontal pass and 1,382 on the slow one, against 8,231 for the whole tree from the same clip. On the neighbouring tree, three of the same four passes returned no trunk frame and no root collar frame at all.

A moving capture does not degrade evenly across the task. It keeps the parts that were never difficult and loses the two views the arborist was hired to look at.

The Same Footage, Read Twice

Everything so far treats the video as a supply of photographs. There is a second reading, and it answers a question the first one cannot.

A phone's built-in GNSS is good to about 3.5 m in the open and worse under a canopy, which is the same order as the spacing between trees in a row. That is enough to find a park and not enough to bind an observation to an individual. Two crews, two visits, one register: unless every observation carries a stable per-tree identity, the second visit cannot be compared with the first, and condition history, the thing an inspection programme exists to produce, never accumulates.

The footage already carries the answer. We ran the project's own clips through a streaming monocular reconstruction, ABot-Recon with the 12-frame local-context model, the same build our abandoned-bicycle work patched for Apple silicon, on a laptop GPU at 0.49 seconds per frame. Frames at 2 fps, no LiDAR, no RTK, no survey control, no GPU server. A ground plane is fitted by RANSAC, and the single unknown of a monocular reconstruction, its scale, is fixed by the one thing we know about the rig, which is how high the camera sits. Points in a band above the ground are then clustered in plan view into trunks.

Sixty seconds of driving along a cherry embankment, filmed from a camera on the car's bonnet, reconstructs like this:

Sixty seconds of driving, reconstructed: elevation along the drive and the plan view with eight individuated trunks

178 metres of embankment, eight trunks individuated with positions, a median spacing of 9.4 m and a median crown top of 5.8 m, from a car moving at walking to cycling pace, in a minute, on hardware that fits in a bag. In the elevation view the trees stand in a row above a flat ground plane, and the canopy the photo diagnosis works from is the top of that structure.

So a tree's identity does not have to come from a GNSS fix. It can come from its position in a reconstruction, a metre-level place in the site's own frame that the next pass can be aligned to. It is the same trick the bicycle-ledger work used to keep one bicycle equal to one ID across an 82-minute route, it is a better key than a 3.5 m GNSS point, and it comes free with footage the crew is shooting anyway.

Three clips, three motions, three different records, and one that returns no geometry at all

Two limits are specific enough to design against. A walk-through reconstructs well over short distances, but its ground plane bends as drift accumulates: visible over fifty metres, and over a full park round it would need control points, which is what the field work on RTK had anticipated. And the close orbit of a single tree, the capture that is best for the diagnosis, returns no usable geometry at all. It is all tree and almost no ground, so there is nothing to fix the scale against.

That second limit changed the protocol.

The Protocol

The resulting protocol: four stops on a 7 m ring, a ground-to-sky tilt at each, and the four part views the model consumes

One phone. Walk a 7 m ring, stop four times, and at each stop make one slow ground-to-sky tilt while the app switches lenses, fires and checks the frames on the spot. About two minutes per tree, standing, with no framing decisions asked of the operator. Blocked directions may be skipped, as the manual already allows. Street trees get two stops on the same pavement, with the telephoto doing the work across the road width. Trees over 20 m, on a slope or over water get a drone booked for the crown case by case, not as standard equipment.

The crew makes no optical decisions. Every number above has been spent before they arrive.

One rule was added after the reconstruction work: keep the ground in frame. Between stops the phone films at walking pace with the wide lens pointed down the path rather than at the tree. Those frames are useless for diagnosis and they are the only frames that make the pass reconstructable, because they carry the ground plane that sets the scale and they tie one stop to the next. The crew is walking between the stops anyway.

What It Costs and What Is Still Open

Arborist annotation ran to roughly ¥250,000 for thirty trees. A programme that buys accuracy by ordering more annotation does not scale. One that buys it by capturing better, with more angles, correctly sampled and bound to the right individual, scales with the site visits the company is already making.

Identity is the part still open, and there are three ways through at different costs. RTK at the belt gives centimetre-level binding for a known equipment spend. Reconstruction gives a position in the site's own frame for free, good to roughly a metre, bounded by drift and improving with every pass that overlaps an earlier one. Re-identification from the imagery binds a tree by its own appearance on top of that position. The last two are the same capability a robot needs to recognise an asset from a viewpoint it has never used before, which is why we are interested in them well beyond this programme.

What We Took Away

The instinct on a project like this is to treat capture as the boring end and the model as the interesting one. It was the other way round. The manual's rules turned out to be a fossil of a lens nobody carries any more, the most attractive rig in the plan was eliminated by an exposure calculation, and the constraint that decides the programme is not how well anything is recognised but whether the right pixels, with a stable identity attached, ever reach the recogniser.

So capture specification is part of the model. Writing it down as numbers a crew can be handed, rather than as advice, is what makes an inspection programme repeatable, and repeatability is what turns a survey into a record.

That record has two halves. The photographs answer what condition this tree is in. The reconstruction answers which tree, where, how big, and whether anyone has been here before. The same two minutes of walking produces both, but only if the capture was designed for both. Design it for the photographs alone and the geometry is gone, not degraded but gone, and with it the one thing that would have let next year's visit know it was looking at the same tree.