Files
star-map/.github/workflows/data-refresh.yml
SenrokaiandClaude Opus 5.5 4ebe663a84 Carry the Hipparcos parallax errors between weekly refreshes, as the Gaia answers already are
a81dd49 added a query of public.hipparcos_newreduction on the ESA archive, cached as
hipparcos-errors-<hash>.csv, and fetchStars requires it: a failed or short answer throws. The
refresh workflow carries only tools/etl/.cache/gaia-dr3-*.csv between runs, a glob that file does
not match, so every weekly run fetched its 117 955 rows live from the archive the cache exists to
spare, and an ESA outage would have failed the refresh even with every Gaia answer cached. Before
a81dd49 a warm cache meant no request to that archive at all.

The cache step now lists both globs. Its key gains "-hipparcos": actions/cache keys are immutable,
and an entry already saved under the old key would be restored without the new file and never
saved again. Checked with Node's path.matchesGlob against the local cache (gaia-dr3-54fdbc7a,
gaia-dr3-hip-785b92fc and hipparcos-errors-f846b045 match, the archive's own files do not) and by
parsing the workflow with PyYAML. A workflow file has no unit test to fail without the change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 20:46:29 +02:00

107 lines
4.8 KiB
YAML

name: Refresh the catalogues
# Re-runs the ETL against the live archives on a schedule and republishes the site when the data
# actually changed, so the catalogues track NASA without anyone touching the code. The archives
# this reads — HYG, the Exoplanet Archive, JPL Horizons, OpenNGC, Gaia — are all anonymous
# public endpoints; no keys are involved.
on:
schedule:
# Mondays 05:23 UTC. An arbitrary minute rather than :00, which is the busiest minute on
# GitHub's cron fleet and the most likely to be delayed or dropped.
- cron: '23 5 * * 1'
workflow_dispatch:
# Never two refreshes at once, and never cancel one mid-push.
concurrency:
group: data-refresh
cancel-in-progress: false
permissions:
contents: write # push the regenerated catalogues to main
actions: write # dispatch the deploy and CI afterwards — see the final step
jobs:
refresh:
name: Fetch, gate, publish
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v5
with:
# The data lives on main and the push below goes to main, whichever ref the workflow
# file itself ran from.
ref: main
- uses: actions/setup-node@v5
with:
node-version: 22
cache: npm
- run: npm ci
# Gaia DR3 is a frozen release: the same query returns the same bytes (a live re-fetch has
# reproduced stars.bin exactly), so its responses are carried from one run to the next
# rather than re-downloaded every week from an archive that times out under load. So is
# van Leeuwen's 2007 Hipparcos reduction, as frozen and on the same archive, whose parallax
# errors fetchStars requires: left out, it was fetched live every week, and an ESA outage would
# have failed the refresh with every Gaia answer cached. The key follows gaia.ts, where the
# queries are written, so a changed query is fetched afresh.
- uses: actions/cache@v4
with:
path: |
tools/etl/.cache/gaia-dr3-*.csv
tools/etl/.cache/hipparcos-errors-*.csv
key: gaia-dr3-hipparcos-${{ hashFiles('tools/etl/sources/gaia.ts') }}
# Every other source is fetched live on this fresh runner. A failed fetch fails the run by
# design — no refresh is better than a partial one. That includes Gaia on a cold cache, by
# two different paths: its Hipparcos cross-match is required, so an unreachable archive
# fails the run from fetchStars itself, while its main query is skipped when unreachable and
# the merge gate in build.ts then refuses a catalogue it contributed nothing to. An archive
# that answers short rather than not at all is caught in fetchGaiaStars.
- name: Rebuild the datasets
run: npm run etl
- name: Detect a real change
id: diff
run: |
if git diff --quiet -- src/assets/data; then
echo "changed=false" >> "$GITHUB_OUTPUT"
echo "The archives published nothing new — catalogues are byte-identical." >> "$GITHUB_STEP_SUMMARY"
else
echo "changed=true" >> "$GITHUB_OUTPUT"
{ echo "Catalogue changes:"; echo '```'; git diff --stat -- src/assets/data; echo '```'; } >> "$GITHUB_STEP_SUMMARY"
fi
# The same gates CI runs, run here instead: the push below is made with GITHUB_TOKEN, and
# GitHub deliberately fires no workflows for such pushes, so the data must be proven
# before it lands rather than checked after.
- name: Unit tests against the new data
if: steps.diff.outputs.changed == 'true'
run: npm test -- --no-watch
- name: Production build against the new data
if: steps.diff.outputs.changed == 'true'
run: npm run build
- name: Commit to main
if: steps.diff.outputs.changed == 'true'
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add src/assets/data
git commit -m "Refresh the astronomical catalogues" \
-m "Scheduled re-run of the ETL against the live archives. Gated on the unit suite and a production build in this same run, because a GITHUB_TOKEN push triggers no CI of its own."
git push origin HEAD:main
# The recursion guard that keeps the bot push from triggering `push` workflows also keeps
# it from deploying, so the deploy — and a visible CI record on the new commit — are
# dispatched explicitly. Dispatch does go through, unlike push events.
- name: Redeploy the site, and put checks on the commit
if: steps.diff.outputs.changed == 'true'
env:
GH_TOKEN: ${{ github.token }}
run: |
gh workflow run pages.yml --ref main
gh workflow run ci.yml --ref main