Files
star-map/.github/workflows/data-refresh.yml
T
SenrokaiandClaude Opus 5 4eb61ff58e Draw each HYG star at Gaia's distance, and keep the ones Hipparcos misplaced
HYG and Gaia were both cut at 250 pc, each on its own distance. A star
Hipparcos put at 200 pc and Gaia at 300 was kept by the first, never
downloaded from the second, and drawn at 200. That is where 83% of the
9 691 mid-magnitude HYG stars without a Gaia counterpart came from, and at
the median Hipparcos had them a third too close. The mirror case, Hipparcos
outside and Gaia inside, dropped the HYG row and left its Gaia entry
anonymous.

Gaia's own Hipparcos cross-match (hipparcos2_best_neighbour, a fixed DR3
table of 99 525 rows) gives a usable Gaia distance for 97 751 of them.
placementDistancePc keeps a star either survey puts inside the cutoff, and
draws every kept star at the better measurement, inside the cutoff or not.
57 121 HYG stars now sit at Gaia's distance. 6 833 of them are past 250 pc:
Zet Per 230 -> 259 pc, 35 Ori 137 -> 330, 44 Cnc 223 -> 613, and the
farthest, HIP 69445, at 8.7 kpc. 3 666 stars that Hipparcos put outside are
now kept, and 3 656 of them give a Gaia entry its name.

The cross-match is required rather than skipped when unreachable. Without
it, every one of those stars would move back to its Hipparcos distance, and
the published map would flip with the archive's availability. The ESA TAP
answered it with a 500 at first and in 102 s on the next try. So fetches
now retry 5xx and network failures twice, after 30 s and 120 s, in the
fetch every source goes through. The refresh job also carries the Gaia DR3
responses from run to run in the Actions cache: the release is frozen, and
a live re-fetch has already reproduced stars.bin byte for byte.

423 651 stars (+10), 61 168 HYG rows folded into Gaia entries (+3 656),
351 597 unnamed designations (-3 656). 10 886 HYG survivors and 23 unmerged
pairs under an arcsecond, both inside the merge gate's ceilings. The same
1 972 exoplanets have a host; KELT-4 A b and MWC 758 c now sit on their
named star.

The HUD's "Radius" becomes "Survey radius": 250 pc is where Gaia is
surveyed to, and no longer the edge of the map.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi
2026-09-11 19:07:12 +02:00

100 lines
4.2 KiB
YAML

name: Refresh the catalogues
# Re-runs the ETL against the live archives on a schedule and republishes the site when the data
# actually changed, so the catalogues track NASA without anyone touching the code. The archives
# this reads — HYG, the Exoplanet Archive, JPL Horizons, OpenNGC, Gaia — are all anonymous
# public endpoints; no keys are involved.
on:
schedule:
# Mondays 05:23 UTC. An arbitrary minute rather than :00, which is the busiest minute on
# GitHub's cron fleet and the most likely to be delayed or dropped.
- cron: '23 5 * * 1'
workflow_dispatch:
# Never two refreshes at once, and never cancel one mid-push.
concurrency:
group: data-refresh
cancel-in-progress: false
permissions:
contents: write # push the regenerated catalogues to main
actions: write # dispatch the deploy and CI afterwards — see the final step
jobs:
refresh:
name: Fetch, gate, publish
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v5
with:
# The data lives on main and the push below goes to main, whichever ref the workflow
# file itself ran from.
ref: main
- uses: actions/setup-node@v5
with:
node-version: 22
cache: npm
- run: npm ci
# Gaia DR3 is a frozen release: the same query returns the same bytes (a live re-fetch has
# reproduced stars.bin exactly), so its responses are carried from one run to the next
# rather than re-downloaded every week from an archive that times out under load. The key
# follows gaia.ts, where the queries are written, so a changed query is fetched afresh.
- uses: actions/cache@v4
with:
path: tools/etl/.cache/gaia-dr3-*.csv
key: gaia-dr3-${{ hashFiles('tools/etl/sources/gaia.ts') }}
# Every other source is fetched live on this fresh runner. A failed fetch fails the run by
# design — no refresh is better than a partial one. That includes Gaia on a cold cache:
# the ETL skips it when unreachable, and the merge gate in build.ts then refuses a
# catalogue it contributed nothing to.
- name: Rebuild the datasets
run: npm run etl
- name: Detect a real change
id: diff
run: |
if git diff --quiet -- src/assets/data; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "The archives published nothing new — catalogues are byte-identical." >> "$GITHUB_STEP_SUMMARY"
else
echo "changed=true" >> "$GITHUB_OUTPUT"
{ echo "Catalogue changes:"; echo '```'; git diff --stat -- src/assets/data; echo '```'; } >> "$GITHUB_STEP_SUMMARY"
fi
# The same gates CI runs, run here instead: the push below is made with GITHUB_TOKEN, and
# GitHub deliberately fires no workflows for such pushes, so the data must be proven
# before it lands rather than checked after.
- name: Unit tests against the new data
if: steps.diff.outputs.changed == 'true'
run: npm test -- --no-watch
- name: Production build against the new data
if: steps.diff.outputs.changed == 'true'
run: npm run build
- name: Commit to main
if: steps.diff.outputs.changed == 'true'
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add src/assets/data
git commit -m "Refresh the astronomical catalogues" \
-m "Scheduled re-run of the ETL against the live archives. Gated on the unit suite and a production build in this same run, because a GITHUB_TOKEN push triggers no CI of its own."
git push origin HEAD:main
# The recursion guard that keeps the bot push from triggering `push` workflows also keeps
# it from deploying, so the deploy — and a visible CI record on the new commit — are
# dispatched explicitly. Dispatch does go through, unlike push events.
- name: Redeploy the site, and put checks on the commit
if: steps.diff.outputs.changed == 'true'
env:
GH_TOKEN: ${{ github.token }}
run: |
gh workflow run pages.yml --ref main
gh workflow run ci.yml --ref main