HYG and Gaia were both cut at 250 pc, each on its own distance. A star Hipparcos put at 200 pc and Gaia at 300 was kept by the first, never downloaded from the second, and drawn at 200. That is where 83% of the 9 691 mid-magnitude HYG stars without a Gaia counterpart came from, and at the median Hipparcos had them a third too close. The mirror case, Hipparcos outside and Gaia inside, dropped the HYG row and left its Gaia entry anonymous. Gaia's own Hipparcos cross-match (hipparcos2_best_neighbour, a fixed DR3 table of 99 525 rows) gives a usable Gaia distance for 97 751 of them. placementDistancePc keeps a star either survey puts inside the cutoff, and draws every kept star at the better measurement, inside the cutoff or not. 57 121 HYG stars now sit at Gaia's distance. 6 833 of them are past 250 pc: Zet Per 230 -> 259 pc, 35 Ori 137 -> 330, 44 Cnc 223 -> 613, and the farthest, HIP 69445, at 8.7 kpc. 3 666 stars that Hipparcos put outside are now kept, and 3 656 of them give a Gaia entry its name. The cross-match is required rather than skipped when unreachable. Without it, every one of those stars would move back to its Hipparcos distance, and the published map would flip with the archive's availability. The ESA TAP answered it with a 500 at first and in 102 s on the next try. So fetches now retry 5xx and network failures twice, after 30 s and 120 s, in the fetch every source goes through. The refresh job also carries the Gaia DR3 responses from run to run in the Actions cache: the release is frozen, and a live re-fetch has already reproduced stars.bin byte for byte. 423 651 stars (+10), 61 168 HYG rows folded into Gaia entries (+3 656), 351 597 unnamed designations (-3 656). 10 886 HYG survivors and 23 unmerged pairs under an arcsecond, both inside the merge gate's ceilings. The same 1 972 exoplanets have a host; KELT-4 A b and MWC 758 c now sit on their named star. The HUD's "Radius" becomes "Survey radius": 250 pc is where Gaia is surveyed to, and no longer the edge of the map. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi
56 lines
2.4 KiB
TypeScript
56 lines
2.4 KiB
TypeScript
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
|
|
import { dirname, join } from 'node:path';
|
|
|
|
const CACHE_DIR = join(process.cwd(), 'tools', 'etl', '.cache');
|
|
const FORCE_REFRESH = process.env['ETL_FORCE_REFRESH'] === '1';
|
|
|
|
/**
|
|
* Downloads `url` as text, caching the raw response under `tools/etl/.cache/<cacheKey>` so
|
|
* re-running the ETL doesn't hit live NASA/astronomy endpoints unless the cache is missing
|
|
* or `ETL_FORCE_REFRESH=1` is set. Keeps the pipeline idempotent and resilient to rate limits.
|
|
*/
|
|
export async function fetchTextCached(url: string, cacheKey: string): Promise<string> {
|
|
const cachePath = join(CACHE_DIR, cacheKey);
|
|
|
|
if (!FORCE_REFRESH && existsSync(cachePath)) {
|
|
return readFileSync(cachePath, 'utf-8');
|
|
}
|
|
|
|
console.log(` fetching ${url}`);
|
|
const text = await fetchText(url);
|
|
|
|
mkdirSync(dirname(cachePath), { recursive: true });
|
|
writeFileSync(cachePath, text, 'utf-8');
|
|
return text;
|
|
}
|
|
|
|
/**
|
|
* How long to wait before each retry of a failed request. The archives this reads are public
|
|
* services that time out under load — the Gaia TAP has answered a five-row join in two and a
|
|
* half minutes and a full one with a 500 — and a weekly refresh that gives up on the first of
|
|
* those publishes nothing that week.
|
|
*/
|
|
const RETRY_DELAYS_MS = [30_000, 120_000];
|
|
|
|
async function fetchText(url: string): Promise<string> {
|
|
for (let attempt = 0; ; attempt++) {
|
|
const response = await fetch(url).catch((error: unknown) => (error instanceof Error ? error : new Error(String(error))));
|
|
if (!(response instanceof Error) && response.ok) {
|
|
return response.text();
|
|
}
|
|
const reason = response instanceof Error ? response.message : `${response.status} ${response.statusText}`;
|
|
// A 4xx is the request's own fault, and waiting will not change the answer.
|
|
const retryable = response instanceof Error || response.status >= 500;
|
|
if (!retryable || attempt >= RETRY_DELAYS_MS.length) {
|
|
throw new Error(`Failed to fetch ${url}: ${reason}`);
|
|
}
|
|
console.log(` ${reason}; trying again in ${RETRY_DELAYS_MS[attempt] / 1000} s`);
|
|
await new Promise((resolve) => setTimeout(resolve, RETRY_DELAYS_MS[attempt]));
|
|
}
|
|
}
|
|
|
|
/** Convenience wrapper around {@link fetchTextCached} that parses the cached response as JSON. */
|
|
export async function fetchJsonCached<T>(url: string, cacheKey: string): Promise<T> {
|
|
return JSON.parse(await fetchTextCached(url, cacheKey)) as T;
|
|
}
|