4 Commits
Author SHA1 Message Date
SenrokaiandClaude Opus 5 44f6a8d086 Refuse a Gaia answer that came back short, and read the body inside the retry
The merge gate asks whether Gaia contributed any stars, never how many. The TAP service truncates
on its own timeout and still serves a well-formed CSV with a 200, ordered by magnitude — so a half
answer is the bright half, which is the half HYG overlaps. Every gate passes: Gaia stars are
present, HYG survivors go down rather than up, unmerged twins can only fall. The weekly job would
publish a catalogue missing two hundred thousand stars and the runner would cache it for the weeks
after. `fetchGaiaStars` now refuses fewer than 95% of the 412 765 rows its query holds, as its
sibling query already did, and refuses an answer that fills the row limit.

`fetchText` retried the request but not the body: a connection reset part-way through the 57 MB
CSV rejected out of the loop, with no wait and no second attempt. The read now happens inside it.

Also corrected: the merge gate's account of the HYG survivors (two thirds of them are stars Gaia
measures but the main query never downloads, since Gaia puts them past the 250 pc cutoff), and the
refresh workflow's comment on what happens when the archive is unreachable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi
2026-09-18 12:12:49 +02:00
SenrokaiandClaude Opus 5 4eb61ff58e Draw each HYG star at Gaia's distance, and keep the ones Hipparcos misplaced
HYG and Gaia were both cut at 250 pc, each on its own distance. A star
Hipparcos put at 200 pc and Gaia at 300 was kept by the first, never
downloaded from the second, and drawn at 200. That is where 83% of the
9 691 mid-magnitude HYG stars without a Gaia counterpart came from, and at
the median Hipparcos had them a third too close. The mirror case, Hipparcos
outside and Gaia inside, dropped the HYG row and left its Gaia entry
anonymous.

Gaia's own Hipparcos cross-match (hipparcos2_best_neighbour, a fixed DR3
table of 99 525 rows) gives a usable Gaia distance for 97 751 of them.
placementDistancePc keeps a star either survey puts inside the cutoff, and
draws every kept star at the better measurement, inside the cutoff or not.
57 121 HYG stars now sit at Gaia's distance. 6 833 of them are past 250 pc:
Zet Per 230 -> 259 pc, 35 Ori 137 -> 330, 44 Cnc 223 -> 613, and the
farthest, HIP 69445, at 8.7 kpc. 3 666 stars that Hipparcos put outside are
now kept, and 3 656 of them give a Gaia entry its name.

The cross-match is required rather than skipped when unreachable. Without
it, every one of those stars would move back to its Hipparcos distance, and
the published map would flip with the archive's availability. The ESA TAP
answered it with a 500 at first and in 102 s on the next try. So fetches
now retry 5xx and network failures twice, after 30 s and 120 s, in the
fetch every source goes through. The refresh job also carries the Gaia DR3
responses from run to run in the Actions cache: the release is frozen, and
a live re-fetch has already reproduced stars.bin byte for byte.

423 651 stars (+10), 61 168 HYG rows folded into Gaia entries (+3 656),
351 597 unnamed designations (-3 656). 10 886 HYG survivors and 23 unmerged
pairs under an arcsecond, both inside the merge gate's ceilings. The same
1 972 exoplanets have a host; KELT-4 A b and MWC 758 c now sit on their
named star.

The HUD's "Radius" becomes "Survey radius": 250 pc is where Gaia is
surveyed to, and no longer the edge of the map.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi
2026-09-11 19:07:12 +02:00
SenrokaiandClaude Opus 5 29d3ddb6ef Say so when Gaia is missing, rather than as 68 000 unmatched stars
With the merge gate in place, a Gaia DR3 outage no longer ships a HYG-only
catalogue: the ETL skips the unreachable source, and validateMerge then fails
on the survivor count. That is the right outcome and the wrong message: "68 000
HYG stars found no Gaia counterpart" sends the reader looking at the merge.
Gaia contributing nothing is now checked first, by name.

Two comments said Gaia was best-effort, in data-refresh.yml and on the merge
in fetchStars. They now say what happens instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi
2026-09-11 18:41:09 +02:00
Claude 8e55964b14 Refresh the catalogues on a schedule, and republish when they change
Mondays 05:23 UTC: re-run the ETL cold against the live archives, and if the
output differs by a byte, gate it on the unit suite and a production build,
commit it to main, and dispatch the Pages deploy. If nothing changed, say so
in the run summary and touch nothing.

The gates run inside this workflow because they cannot run after it: a push
made with GITHUB_TOKEN fires no push workflows at all — GitHub's recursion
guard — so an unguarded push would deploy nothing and be checked by nothing.
The same guard is why the deploy and a visible CI record are dispatched
explicitly afterwards; dispatch events do go through where push events do
not. ci.yml gains a workflow_dispatch trigger for exactly that call.

Also corrects pages.yml's claim that configure-pages enables Pages on first
run. It cannot: the action's `enablement` input requires an admin-scoped
token, which GITHUB_TOKEN is not. If the site has never been enabled, the
first deploy fails at that step and the one-time fix is Settings → Pages →
Source: GitHub Actions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G
2026-08-17 14:05:42 +00:00