Gaia DR3 gives positions for J2016.0, HYG for 2000.0, and the merge matched
them on the sky to one arcsecond without propagating any proper motion.
Sixteen years of motion is 62" for Proxima and 166" for Barnard's Star, so
every star faster than ~62 mas/yr — most of the nearest ones — was kept twice,
some 23 000 in all. The slow ones were matched, and lost: the merge kept
Gaia's row whole, so 102 proper names, 1 336 Bayer/Flamsteed names and
32 000 spectral types became "Gaia DR3 <id>" and "Unknown", and 92 named
exoplanet hosts handed their planets to their anonymous twin.
Gaia is now asked for its proper motions and carried back to J2000 before it
leaves the fetcher. HYG is placed from its own x/y/z columns, which are right
where its `ra` is not: that column was carried from the Hipparcos epoch
without the cos δ its motion needs, 17.9" off for Proxima. A match combines
the two entries — Gaia's position, HYG's name, type, magnitude, colour and id
— instead of choosing one. The tolerance is 15" with a five-magnitude guard,
both set by measurement: 55 457 pairs sit under 1" once the epochs agree, the
Gliese-only entries up to 12" (Ross 248), shifting every entry a quarter of
a degree finds 16 chance neighbours at 15", and the guard keeps Sirius out
of Sirius B's entry. Entries of one source are never merged with each other:
the 1 411 Gaia doubles resolved under 1" are two stars, not one.
Regenerated: 425 071 stars (was 447 410), 56 082 of them Gaia positions
carrying HYG identities; no HYG id or name lost; the sixteen stars nearest
the Sun carry no survey designation; 196 residual doubles, all components
17" or more from their counterpart. Five planets of four bright giants
(7 CMa, HD 81688, omi UMa, xi Aql) lose their host link: their Gaia distance
sits 0.7–1.1 pc from the archive's Hipparcos-based one, past the 0.5 pc the
host match allows. Matching hosts on the sky rather than in space, as the
merge does, is the follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QL6F9Bgfh8SgAiAAcPB9Hw
Scheduled re-run of the ETL against the live archives. Gated on the unit suite and a production build in this same run, because a GITHUB_TOKEN push triggers no CI of its own.
The map held 8750 stars within 50 pc and rendered 371 systems. Both were
lower than they needed to be, for different reasons.
The star catalogue was capped by its own encoding as much as by the
cutoff: one JSON object per star, eight key names repeated each time, 157
bytes a star. At the range HYG actually reaches that is 17 MB to download
and parse before the first frame. So the numbers move into two binary
column stores — positions in stars.bin, which the GPU is handed verbatim,
and id/magnitude/colour/spectral index in stars-meta.bin — and the JSON
keeps only the strings, with 2600 distinct spectral classifications
collapsed to a dictionary. The layout is defined once, in star-catalog.ts,
and the ETL and the app both use it, so the writer and the reader cannot
drift.
The cutoff then goes to 250 pc: 68388 stars, 7.8x as many for 1.7x the
bytes. That is where HYG's measurements stop rather than a round number —
98.6% of its rows are Hipparcos, whose parallaxes are good to about a
milliarcsecond, so beyond 250 pc it would be plotting noise.
Drawing all of them is a separate question from knowing them, and it is
answered separately. The field draws a budget: every star inside 25 pc,
because the nearest are faint red dwarfs and Proxima Centauri is magnitude
11, then the brightest of everything beyond. Search, navigation and the
planet cross-reference still see the whole catalogue. A real GPU would
draw all 68388 without noticing; the budget is for the machines that would
not, and it is one constant.
Systems were limited by something else entirely. The archive data already
shipped named 4735 host stars and only 388 resolved, because the rest lay
outside a 50 pc catalogue — and the cross-reference kept only its own
result, so redoing it meant re-downloading an archive that is not
reachable from here. Host coordinates are now stored with each planet, and
the match is re-resolved at build time against whatever catalogue the run
produced. Even name matching alone, which needs no coordinates and so
works on the records already shipped, rescues 335 planets across 238
systems: 371 renderable systems become 609.
Two selection rules were tuned for a 50 pc bubble and no longer fit.
Tethers followed the Sun's nearest neighbours, which are a speck at this
range, and now follow the brightest; labels were ranked by proximity,
which named whatever sat nearest the middle of the screen, and are now
ranked by brightness — so the view names Canopus, Achernar and Spica
rather than a clump of catalogue designations.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G