Point two cameras at the same street corner and you have created a small, stubborn philosophy problem dressed up as an engineering one. A red truck rolls through the intersection. Camera A, watching from the north, sees a red truck and draws a box around it. Camera B, watching from the east, sees a red truck and draws a box around it. Two boxes, two detections, both correct — and if the map simply plots what each camera reports, it now shows two trucks on a corner that holds one. Nobody lied. The camera did its job; the detector did its job. The map is still wrong.
This is the failure that doesn’t show up in a demo, because a demo has one camera. It shows up the day the network gets dense enough to be useful — the day two contributors happen to be watching feeds whose fields of view overlap, which on any busy corner worth mapping is most of the time. The coverage you wanted and the double-counting you didn’t are the same fact seen from two sides. You cannot have overlapping cameras without overlapping claims, and a live map is a machine for turning claims into dots.
A detector answers “is there a car here?” A map has to answer a harder question the detector never touches: “is this the same car I already have?” Counting isn’t addition. It’s deciding what’s one thing.
The map lies loudest where it matters most
It would be one thing if the error were uniform — a little noise you could shrug off. It isn’t, and the shape of the error is what makes it dangerous. Overlap concentrates where cameras concentrate, and cameras concentrate where there’s something to watch. The sleepy residential block with one feed on it counts honestly. The six-lane downtown junction with four contributors all pointed at rush hour is where the phantoms breed. The map inflates its numbers precisely in the places whose numbers people actually care about.
So the double-counting isn’t a cosmetic blemish; it’s a bias aimed straight at the map’s whole purpose. A live map of a city is, in the end, a way of asking where is it busy? — and a naive one answers by handing back the busiest places exaggerated and the quiet ones faithful. It doesn’t add a little static everywhere. It manufactures traffic exactly where it’s already crowded, then reports the crowd it invented. A map you can’t trust in the busy places is a map you can’t trust, because the busy places are the entire reason you opened it.
Fusion is a question about identity, not vision
The instinct is to fix this upstream, with better seeing — a smarter detector, cleaner boxes, tighter confidence. But no amount of vision helps here, because both detections are right. There genuinely is a red truck in Camera A’s frame and a red truck in Camera B’s frame. The question fusion has to settle isn’t “is this a truck?” It’s “are these two trucks the same truck?” — and that’s not a perception problem at all. It’s an identity problem, the same one a newsroom faces when two reporters phone in the same story: not is each account true but is it one event or two?
Thyseus answers it the way you’d answer it about the reporters — by asking whether the two accounts can be reconciled into one. Once each detection is projected out of pixels and onto the ground, a detection stops being a box in a frame and becomes a claim about a place: a truck, here, at this second. Two claims that land on the same patch of asphalt within the same slice of time are almost certainly one truck seen twice, not two trucks improbably parked on top of each other. Fusion merges them into a single track and lets the map count that track once. Two reports go in; one thing comes out.
The care is all in the word almost. Merge too eagerly and two cars stopped bumper-to-bumper at a light become one — you’ve erased a real vehicle to tidy up a phantom. Merge too timidly and every overlap leaks a ghost, and you’re back to inflating the busy corners. The tolerance — how close in space, how close in time before two claims collapse into one — is the whole craft, and it can’t be a global constant. A crowded intersection where cars sit inches apart needs a tighter rule than an empty highway shoulder where the nearest two things are a truck and a road sign. Same as the map’s argument about when: the honest answer is a judgment the system makes out loud, not a threshold someone hard-coded and forgot.
Counting is the part that has to be right
It’s tempting to file all of this under polish — the map basically works, this just cleans up the edges. It’s the reverse. Detection is the part that looks like the product in a screenshot; fusion is the part that decides whether the product is true. A map that spots every car but can’t tell one car from its own echo isn’t a map with a rough edge. It’s a hall of mirrors that happens to render nicely.
And it compounds. Every number the map is ever asked for — how busy is this corner, how did traffic here change this hour, is this junction worse on Fridays — is a count, and a count is only ever as honest as the system’s grip on what counts as one thing. Get fusion right and every downstream number inherits the honesty for free. Get it wrong and there’s no analytics clever enough downstream to recover the truth, because the mistake was baked in at the moment of counting, before any query was ever asked.
That’s why Thyseus treats “is this the same car?” as a first-class question and not an afterthought. The glamorous half of a live map is the seeing — the boxes lighting up on a dozen feeds at once. The half that makes it a map you can believe is the quiet one, running every second on the overlap: two cameras saw one car, and the map has to know it was one.