Every sampled frame of v3 and v4 on one map

Two-dimensional UMAP embedding of SigLIP features for 24,254 sampled frames: 15,040 from v0.4.0 (filled markers) and 9,214 from v0.3.0 (hollow markers), coloured by title. Nearby points are visually similar frames. The 3,032 larger markers, eight per session, show the recorded frame on hover and, for v0.4.0, the engine depth map of the same frame. Clicking a title in the legend hides or shows it. The two numbers after each title are its leave-one-out Vendi contributions in appearance and in geometry: the loss in effective number of distinct frames when the title is removed, divided by the loss when the same number of frames is removed at random (1.0 = equal to random footage of the same size; kernel exp(-d/0.10)).

what the camera saw
what the engine measured

Positions come from a neighbourhood-preserving projection of the 768-number frame fingerprints; distances between clusters carry no meaning, neighbourhoods do. Hover follows the nearest frame-carrying dot only.