83dc46416b98616d79b00e60032b745291c557d4
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1a5785fb63 |
Keep the row cap live for the queries that can reach it
Gating it on `jobsQuery` switched it off for every override that *widens* the query — which is the only way to fill `select top N` at all. `ETL_GAIA_MAGNITUDE_LIMIT=14` asks for 500 000 rows, the sky holds more, and the answer is the limit rather than the filters: exactly what the tripwire is for, and it no longer fired. It now reads the row limit itself, so only a deliberately smaller slice is silent. Measured with a synthetic answer of exactly 500 000 rows in the cache, under the key the widened query hashes to: refused. With the `jobsQuery` gate back, the same run keeps 500 000 Gaia stars and goes on to publish them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
b7f277ea04 |
Answer the review: a short answer makes more survivors, and must not be skipped
Three things this got wrong. The direction: truncating Gaia leaves the HYG rows whose counterpart it dropped without one, so survivors rise — 10 886 today, 12 711 at half the rows, 16 258 at a third — which the comment claimed was the other way, and which decides whether the 15 000 ceiling can be leaned on at all (it catches a truncation past about two thirds, and nothing shallower). The throw: `fetchStars` catches everything a source throws and skips it, so a truncated CSV was reported as "the archive was unreachable" one step after `writeStarAssets` had already overwritten the published catalogue. Marked with `GaiaAnswerError` and rethrown there, so an answer that cannot be worked with fails the run where it happened. Measured end to end in a throwaway working directory, 300 000 rows in the cache: fails, names the cache file to delete, assets untouched. With the rethrow taken back out again: assets written, then "the archive was unreachable". The row limit: `rows.length >= ROW_LIMIT` is true for every reduced ETL_GAIA_ROW_LIMIT, so the tripwire fired on exactly the deliberate slice the override exists for — and told the operator to raise it. Gated on the same flag as its neighbour. `ETL_GAIA_ROW_LIMIT=20000` now runs through; without the gate it dies on the limit it was given. Also: the row floor names the one cache file it is about rather than a glob that takes the Hipparcos cross-match with it, and says an edited query is a third reason it can fire — DEFAULT_QUERY_ROWS now sits under the query it counts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
44f6a8d086 |
Refuse a Gaia answer that came back short, and read the body inside the retry
The merge gate asks whether Gaia contributed any stars, never how many. The TAP service truncates on its own timeout and still serves a well-formed CSV with a 200, ordered by magnitude — so a half answer is the bright half, which is the half HYG overlaps. Every gate passes: Gaia stars are present, HYG survivors go down rather than up, unmerged twins can only fall. The weekly job would publish a catalogue missing two hundred thousand stars and the runner would cache it for the weeks after. `fetchGaiaStars` now refuses fewer than 95% of the 412 765 rows its query holds, as its sibling query already did, and refuses an answer that fills the row limit. `fetchText` retried the request but not the body: a connection reset part-way through the 57 MB CSV rejected out of the loop, with no wait and no second attempt. The read now happens inside it. Also corrected: the merge gate's account of the HYG survivors (two thirds of them are stars Gaia measures but the main query never downloads, since Gaia puts them past the 250 pc cutoff), and the refresh workflow's comment on what happens when the archive is unreachable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
4eb61ff58e |
Draw each HYG star at Gaia's distance, and keep the ones Hipparcos misplaced
HYG and Gaia were both cut at 250 pc, each on its own distance. A star Hipparcos put at 200 pc and Gaia at 300 was kept by the first, never downloaded from the second, and drawn at 200. That is where 83% of the 9 691 mid-magnitude HYG stars without a Gaia counterpart came from, and at the median Hipparcos had them a third too close. The mirror case, Hipparcos outside and Gaia inside, dropped the HYG row and left its Gaia entry anonymous. Gaia's own Hipparcos cross-match (hipparcos2_best_neighbour, a fixed DR3 table of 99 525 rows) gives a usable Gaia distance for 97 751 of them. placementDistancePc keeps a star either survey puts inside the cutoff, and draws every kept star at the better measurement, inside the cutoff or not. 57 121 HYG stars now sit at Gaia's distance. 6 833 of them are past 250 pc: Zet Per 230 -> 259 pc, 35 Ori 137 -> 330, 44 Cnc 223 -> 613, and the farthest, HIP 69445, at 8.7 kpc. 3 666 stars that Hipparcos put outside are now kept, and 3 656 of them give a Gaia entry its name. The cross-match is required rather than skipped when unreachable. Without it, every one of those stars would move back to its Hipparcos distance, and the published map would flip with the archive's availability. The ESA TAP answered it with a 500 at first and in 102 s on the next try. So fetches now retry 5xx and network failures twice, after 30 s and 120 s, in the fetch every source goes through. The refresh job also carries the Gaia DR3 responses from run to run in the Actions cache: the release is frozen, and a live re-fetch has already reproduced stars.bin byte for byte. 423 651 stars (+10), 61 168 HYG rows folded into Gaia entries (+3 656), 351 597 unnamed designations (-3 656). 10 886 HYG survivors and 23 unmerged pairs under an arcsecond, both inside the merge gate's ceilings. The same 1 972 exoplanets have a host; KELT-4 A b and MWC 758 c now sit on their named star. The HUD's "Radius" becomes "Survey radius": 250 pc is where Gaia is surveyed to, and no longer the edge of the map. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
037545d036 |
Answer the review: a name two stars answer to names neither, and NaN is not a proper motion
Three guards the matcher was missing, none of which changes a byte of the regenerated data — the ETL re-run after them is identical — and all three now have a test that fails without them. A proper motion that is not a number poisoned every comparison rather than one: NaN loses every `<` it appears in, so `cosine < minCosine` was false for every star, each one reached the distance guard, and the last one in catalogue order won — a confident wrong answer, order-dependent, where the honest answer is "no match". The archive's own parser never produces one (parseOptionalNumber maps a blank cell to undefined), but the matcher is exported for offline re-cross-referencing and a caller reaching for bare Number() is exactly the coercion the CSV helper documents as having caused two prior bugs. An unusable motion now reads as no motion. Normalizing a name strips the dot, so `Gl 55.2` and `Gl 552` — two stars 135 degrees apart — share one key, and the index kept whichever came last; 64 such groups exist in the catalogue, among them `Gl 84.1A`/`Gl 841A` and `HD 96600` twice. A name that names two stars names neither, so ambiguous keys are dropped and the query goes to the sky, where direction settles it. No archive hostname lands on one today, which is why the data is unchanged. And the cache is keyed by the whole request rather than the query alone, here and in gaia.ts: fetchTextCached records only that some response arrived, so an endpoint edit would have kept serving the old host's bytes — the same silent staleness the query hash was added to close. The tests now discriminate what the comments claim. Eight mutants, each caught: judging only the published position, only the carried-back one, judging each star on its worse epoch rather than its better, letting a distance-rejected star claim best-so-far and shadow the true host behind it, an unguarded proper motion, a last-wins name index, a fixed angular tolerance instead of a transverse one, and no distance guard at all. The GJ 887 test grew a decoy standing halfway along the star's own track: it is nearer than Lacaille 9352 at the published position and nearer at the worse of the two epochs, so it wins unless both epochs are tried and the better one decides — the property the test's comment had been claiming untested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
dc2ce08694 |
Answer the review: direction settles distance, brightness is one-sided, and a lost id stops the scene
Three findings from the adversarial review of the merge, all reproduced. The distance test was hiding 1 489 stars that sit under an arcsecond from their Gaia entry with a Hipparcos parallax off by half — thirty of them at a false few parsecs from the Sun (HIP 82724 at 3.7 pc, where Gaia has it at 62.8) — and the first audit did not see them because it counted residual doubles through the same 50 % filter. Under three arcseconds the distances are now not consulted: a coincidence of direction that close is never chance at this depth (the quarter-degree shift finds none), and the parallax is the thing to fix. Brightness keeps its say at any separation, and is now one-sided: a folded entry may be five magnitudes fainter (a red dwarf in V against G) but not one brighter, because an entry a magnitude brighter than what is already at that spot is a primary Gaia does not carry — Almach, Alfirk and Ashlesha had all been folded into their companions' entries, 93 in all. The sky grid wraps at 0h. The Gaia query orders by source_id after G, so the row order — and the ids assigned from it — is a function of the archive's content rather than of the server's plan for 20 064 ties; the cache key is a hash of the query. And a bookmark to a star id the catalogue no longer holds — 56 000 Gaia ids change with this — sent the scene through reconcileSelection, enterSystem, its decline, finishTransition and reconcileSelection again until the stack overflowed. The selection is cleared instead, at the one place every path goes through. Regenerated: 423 641 stars, 57 512 HYG identities on Gaia positions, no HYG id or name lost, no star within 20 pc left with an unclaimed Gaia entry under an arcsecond. 403 HYG survivors still have an unclaimed Gaia entry within 60": 13 under an arcsecond, where the brightness guard does not trust HYG's magnitude, and the rest components 3" to 60" from their counterpart. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QL6F9Bgfh8SgAiAAcPB9Hw |
||
|
|
08534279fb |
Bring Gaia to HYG's epoch before merging, and keep a star's name when it matches
Gaia DR3 gives positions for J2016.0, HYG for 2000.0, and the merge matched them on the sky to one arcsecond without propagating any proper motion. Sixteen years of motion is 62" for Proxima and 166" for Barnard's Star, so every star faster than ~62 mas/yr — most of the nearest ones — was kept twice, some 23 000 in all. The slow ones were matched, and lost: the merge kept Gaia's row whole, so 102 proper names, 1 336 Bayer/Flamsteed names and 32 000 spectral types became "Gaia DR3 <id>" and "Unknown", and 92 named exoplanet hosts handed their planets to their anonymous twin. Gaia is now asked for its proper motions and carried back to J2000 before it leaves the fetcher. HYG is placed from its own x/y/z columns, which are right where its `ra` is not: that column was carried from the Hipparcos epoch without the cos δ its motion needs, 17.9" off for Proxima. A match combines the two entries — Gaia's position, HYG's name, type, magnitude, colour and id — instead of choosing one. The tolerance is 15" with a five-magnitude guard, both set by measurement: 55 457 pairs sit under 1" once the epochs agree, the Gliese-only entries up to 12" (Ross 248), shifting every entry a quarter of a degree finds 16 chance neighbours at 15", and the guard keeps Sirius out of Sirius B's entry. Entries of one source are never merged with each other: the 1 411 Gaia doubles resolved under 1" are two stars, not one. Regenerated: 425 071 stars (was 447 410), 56 082 of them Gaia positions carrying HYG identities; no HYG id or name lost; the sixteen stars nearest the Sun carry no survey designation; 196 residual doubles, all components 17" or more from their counterpart. Five planets of four bright giants (7 CMa, HD 81688, omi UMa, xi Aql) lose their host link: their Gaia distance sits 0.7–1.1 pc from the archive's Hipparcos-based one, past the 0.5 pc the host match allows. Matching hosts on the sky rather than in space, as the merge does, is the follow-up. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QL6F9Bgfh8SgAiAAcPB9Hw |
||
|
|
efa9e4084a |
Draw the whole catalogue, and build the aggregation the rest would need
Two things, one verified and one that cannot be. The render budget is now the whole catalogue: 68388 stars, one instanced draw call, which is what a GPU should be asked to do. The budget itself stays, because the catalogue is meant to grow past what any machine should draw at once — Gaia alone could contribute a million — and at that point the selection is what keeps the field legible rather than a grey wash. A `?stars=` override handles the machines that cannot, including the software rasterizer the end-to-end suite runs against, whose frame rate is two orders of magnitude below a real GPU's and which was measuring the rasterizer rather than the app. The aggregation is the second thing, and none of it has run. Every ESA, NOIRLab, SDSS and Euclid endpoint is unreachable from here — only GitHub raw is, which is why HYG and OpenNGC are the current sources. So this is infrastructure and a Gaia query written against the published DR3 schema, not data. What the framework encodes is that these surveys are not interchangeable. The distinction is not size but whether a catalogue knows how far away its objects are, because a 3D map cannot place a star it only has a direction for. Gaia is the only one of the five that can add stars here, because it is the only one that measures parallaxes. DECaPS2 has fifty times Gaia's object count and photometry alone — not one of its 3.32 billion objects can be placed in depth. Euclid's bulge is 8 kpc away, where a parallax is microarcseconds; its contribution would be imagery. SDSS-V and SAGA are keyed to stars something else already places, so they enrich rather than extend. Those roles are recorded as data the ETL prints, not as prose that can drift. Overlapping catalogues are reconciled on direction rather than on 3D proximity, which is the one non-obvious part. Two surveys agree on a star's direction to within an arcsecond and disagree on its distance by tens of per cent, so a star at 200 pc is 50 pc from itself between catalogues while being unmistakably the same object. Matching in 3D would need a tolerance so loose it swallowed real neighbours. The better parallax wins where both reach; where only one does, the star stays. Names become dense-with-holes with a source dictionary, because a survey catalogue has no proper names — writing "Gaia DR3 4472832130942575872" once per star would cost 25 MB per million to repeat what two adjacent fields already say. An empty entry costs three bytes and is regenerated on load. The Sun needed its own case in the merge: it sits at the origin, has no direction to compare, and appears in every catalogue. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G |