5fb0d45623314e8a7fab0c8ee9ffe4e5517f619b
20
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
23547defe0 |
Read an archive host's colour off the dwarf sequence at its temperature, and say it was not measured
For an archive-placed host with a temperature and no B magnitude, fetchExoplanets took B−V from Ballesteros' blackbody fit, which runs 0.1 to 0.2 redder than Pecaut & Mamajek's dwarf sequence below 3 800 K, and the luminosity then read its correction off that sequence at that colour: 3 500 K came back as 3 102 K with a correction 1.15 magnitudes too large, anything under about 3 170 K was clamped to B−V 2.00, and the card printed it as a measured "Colour B−V 2.00". CFBDSIR J145829+101343, a 580 K brown dwarf, read "Spectral type ~M6, from colour". temperatureToColorIndex now reads the table itself backwards, interpolating B−V between the two types the temperature falls between, so the temperature and correction read back off the colour are the table's at that temperature; it has no answer outside 2 420 to 31 400 K. The colour is flagged colorFromTemperature, a fifth bit in the photometry byte (the format, README and the ETL's round-trip check follow), and the card prints it "B−V 1.66, from its temperature", marked derived. From cache: 57 archive stars change colour; 54 carry the flag and 3, CFBDSIR J145829+101343 among them, now have none. For the 47 of them the archive gives a luminosity, the one derived from magnitude and colour moves from a median 0.228 dex off it to 0.124; Kepler-445 (3 157 K) from 0.0282 L☉ to 0.0080 against the archive's 0.0079. Its colour goes from 2.00 to 1.67. The card shows the archive's luminosity where it has one since earlier on this branch, so this is the figure used for the rest and for their radii. Controls, each failing its named test: the nearest hotter row taken without interpolating (2 of 803 failed), no refusal outside the table, the flag not encoded, and the card calling the colour measured (1 of 803 each). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
f82b5e1c74 |
Hold the published catalogue to what its host radii, distance errors and Gaia-distance flags come to
Three things the branch added could vanish from the data with every check passing, as the review
showed with guarded ETL mutants: dropping st_rad or st_teff from each planet left 0 of 6 354
carrying their host's radius or temperature, storing Gaia's parallax_error in milliarcseconds
rather than over the parallax changed 431 464 distance readouts, and dropping the flag fetchStars
sets relabelled 8 129 HYG stars placed at Gaia's distance "HYG". The round trip compares decoded
with encoded, and the counts of missing bands and errors do not look at what the values are.
validateExoplanets now requires 90 % of planets to carry their host's radius and temperature
(measured 6 030 and 6 054 of 6 354). validateStars requires the median relative error of the
distances at Gaia's to be under 1 % (measured 0.33 %), at most 1 % of stars to have a parallax
distance with an error of a fifth or more (measured 1 109, 0.24 %; the archive's, which are on the
distance and never ranged, are left out), and at least 6 000 HYG stars flagged at Gaia's distance
(measured 8 129). Figures from an ETL run from cache on this branch.
Each proved by a guarded mutant in a scratch copy of the ETL, run from the cache, failing with its
own message: host radius dropped ("Only 0 of 6354 exoplanets carry their host's radius and 6054 its
temperature"), host temperature dropped ("6030 ... and 0"), the Gaia error left in mas ("The median
error of Gaia's distances is 1.92 %"), the same with the median check waived ("17689 parallax
distances have an error of a fifth or more"), and the flag dropped ("Only 0 HYG stars are flagged
at Gaia's distance"); a no-op edit completed. No ceiling is loosened.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
70cbdae306 |
Drop the "..." Hipparcos leaves on a spectral type, which read as the app cutting it short
2 127 stars, Sirius among them, read "Spectral type A0m...", in search rows, route options and the star card. The mark is Hipparcos's own: every one of the 3 803 HYG rows ending in it has a HIP number, and the catalogue ends a classification it does not print in full with it. The app truncated nothing, but the text reads as if it had. fetchStars now trims a trailing "..." from HYG's spectral types. The dictionary in stars-index.json goes from 3 056 distinct types to 2 883, and from 598 dotted ones to none. The names, sources and source indices are unchanged, and so is every other column of stars-meta.bin bar the type indices. gzip -9 of the index goes from 4 015 865 to 4 015 205 bytes. build.ts validateStars now fails a catalogue with any type ending in "...". Run with the trim removed, the ETL failed on 2 127 such types. Two ETL runs from cache wrote identical stars-meta.bin and stars-index.json. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
a81dd491ad |
Say which band each star was measured in, whose catalogue it is, and how sure its distance is
The readout named Hipparcos, Yale and Gliese while 378 775 of the 455 608 stars are described by Gaia DR3, printed one "Magnitude" for G and V alike, and gave every distance to the parsec. The band cannot be read off the star's source, which records whose position it has: 62 002 stars Gaia places keep HYG's V and B-V, and the 575 hosts renamed after their planets are Gaia's, in G. stars-meta.bin gains two byte columns, 14 to 16 bytes a star (6 378 512 to 7 289 728 bytes; gzip -9 2 804 049 to 3 110 991). One holds the magnitude's band (V 76 555 stars, G 378 744, none 309, whose magnitude is a stand-in), which colour the colour index is (B-V 75 203, BP-RP 376 703), and whether the distance is Gaia's parallax. The other holds the distance's relative error as its square root in 255ths: a step is 0.08 % of distance at 1 %, 0.35 % at 20 %, and 100 % is the top. encode, decode, BYTES_PER_STAR_META and the build.ts round trip cover both; no workflow reads the format. Where the errors come from: - Gaia rows keep the parallax_error their query already fetched: median 0.3 %, 90th percentile 1.2 %, at most 20 %, the query's own cut. - A HYG star at Gaia's distance takes the cross-match's parallax_over_error, and a star merged into a Gaia entry keeps that entry's error with its position. - The 3 067 Hipparcos stars that keep their Hipparcos distance, Rigel, Deneb and Alnilam among them, take e_plx from van Leeuwen's 2007 reduction: a new cached query of public.hipparcos_newreduction on the ESA archive, whose 117 955 rows HYG's distances invert. - The archive's stars take sy_disterr1/2 from pscomppars, in a query and cache file of their own so the composite rows already cached were not refetched. - 439 distances have no published error: 357 Gliese rows and 82 archive hosts. Of the errors, 392 786 are 1 % or less and are not printed; 61 156 print as "117 ± 12 pc" to the distance's own digits; 1 196 between 20 and 100 %, and 30 past it, print as the range the parallax gives, since a symmetric error in parallax is a lopsided one in distance. The star card (measured on the dev server) now reads, for example: - Rigel: 265 ± 23 pc, V 0.18, B-V -0.03, source HYG. - Deneb: 433 ± 60 pc. - Alnilam: "476 pc to 833 pc" (Hipparcos 1.65 ± 0.45 mas). - Gaia DR3 5612323414549657984: 111 ± 2 pc, G 4.63, BP-RP -0.15, source Gaia DR3. - Proxima Centauri: 1.30 pc, V 11.01, source "HYG, Gaia DR3 distance". - TRAPPIST-1: G 15.62, BP-RP 4.90, source Gaia DR3. - Kepler-186: V 15.14, source NASA Exoplanet Archive. The neighbourhood's subtitle reads "Gaia DR3 378,775 · HYG 73,556 · NASA Exoplanet Archive 3,277", counted by the catalogue describing each star. build.ts validateStars now fails a catalogue with more than 1 000 stars without a band (309 today) or without a distance error (439). Dropping G from the Gaia rows gave 379 040 without a band, and dropping their parallax_error gave 441 216 without an error; both runs failed. Decoding the catalogue in Node took a median 29 ms before and 24 ms after (nine runs each, within noise). In the app, five cold boots gave a 654-786 ms long task after the data landed and the HUD at 1.83-2.07 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
4c8e4a02f3 |
Give every planet with a distance a star, from the archive where the catalogue has none
After matching, 4 237 planets still had no star: their hosts are too faint for either Gaia query (fainter than G 12 past 50 pc) or too far (past 250 pc), Kepler-186 at 177.6 pc among them. fetchExoplanets now adds one star per such host from the archive's own figures, 3 277 of them, whenever the archive gives a position and a distance: - position carried back from J2016, where the archive publishes it (741 of the 746 matched hosts moving over 100 mas/yr sit nearer their star carried back, a median 0.11" against 3.47" as published), and placed at sy_dist; - magnitude in V (2 999 hosts), else Gaia G (13), the band the catalogue's Gaia stars are already in; 265 have neither, 128 KMT, 95 OGLE and 32 MOA microlensing hosts at a median 6.2 kpc among them, and take the ETL's faint stand-in of 15; - colour as B-V from the archive's B and V, else from st_teff through a new temperatureToColorIndex (Ballesteros 2012, inverted; the Sun's 5 772 K gives 0.65), else left to st_spectype; - ids from 1 070 000 000, past Gaia's two ranges and under the 2^30 validateStars enforces; source "exoplanet-archive". Hosts past 250 pc are included: 399 of the 3 277 are within 250 pc, 1 962 between 250 pc and 1 kpc, 916 beyond. The drawn budget still chooses what is drawn (70 000 of 455 608). Planets with a star: 2 090 -> 6 327 of 6 354. The other 27 have no distance in either table (Luhman 16 A, mu2 Sco, PSR B1620-26 among them). Systems in the Solar Neighbourhood readout: 1 450 -> 4 736. validateExoplanets now refuses a catalogue where fewer than 99.5 % of planets have a host (measured 99.58 %); dropping the added stars fails it at 2 090, and dropping the composite fill at 6 227. No added star sits within an arcsecond of a catalogue star (validateMerge's twins stay at 23); 5 of the 399 within 250 pc have one within a minute of arc, VHS J125601.92-125723.9 (archive 12.7 pc) 5.8" from a Gaia entry at 21.2 pc the likeliest duplicate. In the app, searching TRAPPIST-1 or Kepler-186 and picking the star enters a system with its seven and five planets drawn. gzip -9: stars.bin 5 030 741 -> 5 067 864 B, stars-meta.bin 2 784 614 -> 2 804 064, stars-index.json 4 003 584 -> 4 015 882, exoplanets.json 342 467 -> 446 532 (host parameters, both commits). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
494fb5616b |
Find the hosts the catalogue already holds, and name them after their planets' star
The archive's default-parameter rows leave sy_dist blank for 100 planets, TRAPPIST-1's seven among them, so the matcher could not place a host the catalogue had drawn since the Gaia nearby query (Gaia DR3 2635476908753563008, 12.47 pc). fetchExoplanets now also reads the Planetary Systems Composite table (pscomppars) and fills a default row's blank host cells from it. Only host columns: a composite row takes each column from its own reference, so orbits still come from the default row alone, one fit per planet. Where both tables give a distance or a position they agree on all 6 225 and 6 352 rows. Three hosts the catalogue holds by name failed the distance ratio test because the archive's sy_dist, from TICv8, contradicts its own parallax: Lalande 21185 (GJ 411) 5.68 pc against 392 mas, Luyten's Star (GJ 273) 5.92 against 263 mas, Struve 2398 B (Gl 725 B) 6.84 against 285 mas. resolveHostStarId now takes the archive's parallax (sy_plx) as a second distance the ratio test accepts. It is a second chance, not a replacement: 47 of 5 959 systems disagree past the tolerance, and for faint far hosts the inverse parallax is the worse figure (K2-238: 538 pc by sy_dist, 6 779 by parallax). Planets on a catalogue star: 2 071 -> 2 090 of 6 354. The 19 gained are TRAPPIST-1 (7), GJ 273 (2), GJ 411 (2), Gl 725 B, HD 62509 (Pollux), K2-65, TOI-2267 A and B, and two brown-dwarf hosts, 2MASS J02192210-3925225 and DENIS-P J082303.1-491201. No planet lost or changed its host. A matched host whose catalogue name is a bare Gaia designation now takes the archive's host name, so search finds TRAPPIST-1, Teegarden's Star, TOI-700, LP 791-18 and K2-18: 575 stars renamed. validateExoplanets refuses a host that is still only a designation, and the star assets are rewritten after the exoplanets for that reason; stars.bin and stars-meta.bin are unchanged. TOI-2267 A and B both land on one Gaia entry, which takes the name TOI-2267 A. Each planet also carries its host's radius, effective temperature and luminosity (10^st_lum), and a mass for 6 344 planets instead of 5 474, from the default row where it gives one and the composite table otherwise, for the star-physics step. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
c01d3ec2bc |
Read OpenNGC's addendum, so the Pleiades and the Large Magellanic Cloud are on the map
OpenNGC keeps the objects no NGC or IC number covers in a second file, addendum.csv, with the same 32 semicolon-separated columns as NGC.csv. The ETL only ever fetched NGC.csv, so the brightest deep-sky object in the sky, the Large Magellanic Cloud (V 0.29), was missing while the Small one was drawn, and so were the Pleiades (M45), the Hyades, the Horsehead and the Coalsack. fetchDeepSky now fetches the addendum beside NGC.csv, caches it as openngc-addendum.csv, and runs its rows through the same loop. Of its 64 rows, 27 pass the existing filters: all 22 named or Messier rows except M40, which OpenNGC types as a double star, and M102, typed as a duplicate of M101; plus seven anonymous open clusters brighter than V 9 (H05, H20, H21, Mel071, Mel101, Mel105, MWSC3171). deepsky.json goes from 463 to 490 objects (galaxies 73 -> 85, clusters 294 -> 306, nebulae 96 -> 99), with no existing record changed. Messier coverage goes from 106 to 107 of 110. 340 of the 490 have a distance: the Pleiades 135.8 pc and the Coma Star Cluster 85.9 pc from their parallaxes; the Hyades and the Local Group dwarfs honestly have none. validateDeepSky now requires the Andromeda Galaxy and the Small Magellanic Cloud from NGC.csv, the Large Magellanic Cloud and the Pleiades from the addendum, and at least 107 Messier objects. Both checks were run against the real catalogue with the addendum removed: the ETL fails on "Deep-sky object ESO056-115 is missing", and with the required list emptied, on "Only 106 Messier objects were produced". In the app, on the dev server: 490 sprites. The Pleiades sprite lies 0.00182 deg from Alcyone (0.00183 deg from the published coordinates); the LMC 18.4209 deg from Canopus and 26.8215 deg from Achernar (18.4209 and 26.8215 published); the Hyades 2.2591 deg from Aldebaran (2.2591). The twelve labelled deep-sky objects now open with the Large Magellanic Cloud and the Pleiades and include Brocchi's Cluster, in place of h Persei, chi Persei and NGC 2516. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
74b93a0428 |
Fill in the Sun's faint neighbours from Gaia, and fold the Gliese entries they repeat
Gaia's G < 12 cut took a quarter of what lies within 10 pc: the red, brown and white dwarfs most of the neighbourhood is made of, which the map only had where Gliese happened to list them. Teegarden's Star was missing, and with it its three planets' host. A second query fetches the complement out to 50 pc: G >= 12 or no G, parallax > 20 mas, with pmra/pmdec so the J2016 -> J2000 propagation applies. The quality filter was chosen by counting. Parallax over error > 5 keeps 39 751 of the 39 764 sources that pass the floor, but the Gaia Catalogue of Nearby Stars (GCNS; Gaia Collaboration, Smart et al. 2021) rejects 11 752 of them as spurious: median G 20.2, astrometric excess noise 5.4 mas against 0.15 for the ones it keeps, 6 933 toward the Galactic centre. So the query joins the GCNS main table (EDR3 astrometry and source ids, which DR3 carries unchanged) and keeps 28 012 rows; the error cut stays and costs no GCNS source, brown dwarfs included. The query has its own row-count floor (28 012) and the row-limit cap, and is ordered by (phot_g_mean_mag, source_id); a second ETL run reproduced stars.bin, stars-meta.bin and stars-index.json byte for byte. Its ids start at 1 050 000 000, clear of the main query's and under 2^30, which V8 keeps unboxed: numbered from 2 000 000 000 they made the app's boot task 230 ms longer (medians of five interleaved runs, 1.41 s against 1.18). validateStars now refuses an id outside 0 to 2^30. The Gliese entries these stars duplicate were not folded: HYG carries them with positions off by up to a minute of arc and photometric distances, so they missed the 15" tolerance or failed the distance test. Their proper motions, which Gliese measured well, give them away: isSameStar now takes two entries moving within 20 % of each other as one star up to 60" apart, whatever their distances, brightness still permitting. Of the 602 Gliese-only rows left without a counterpart, 253 have such a Gaia entry; with every entry shifted a quarter degree, none does. fetchStars passes HYG's motions only for rows without Hipparcos astrometry: given them too, 15 Hipparcos stars took a co-moving companion's Gaia entry and the cross-catalogue pairs under an arcsecond went from 23 to 35. Of HYG's 1 200 stars fainter than V 12 within 25 pc, 1 024 now sit on a Gaia position (325 before); of the 176 left alone, 43 still have a Gaia entry 3-60" away (218 without the motion rule), some of them real companions. Measured on the rebuilt catalogue, against the GCNS (sources with parallax > 100, 40 and 20 mas): within 10 pc 336 -> 372 (GCNS 312; the map adds 60 HYG-only stars) within 25 pc 3 652 -> 5 560 (GCNS 5 111) within 50 pc 13 702 -> 40 916 (GCNS 40 231) 452 331 stars (+27 214). Proxima, Barnard's Star, Wolf 359, Rigel, Deneb and Alnilam are all present by name; Teegarden's Star is Gaia DR3 35227046884571776 at 3.83 pc and hosts its three planets. Luhman 16 is not in Gaia DR3 with a parallax (5353626573555863424 has a two-parameter solution) and stays absent. 94 more exoplanets find a host (2 071), none changes host. TRAPPIST-1 is now drawn (Gaia DR3 2635476908753563008, 12.47 pc) but its planets are not yet matched to it. Merge gate, ceilings unchanged: HYG rows without a Gaia counterpart 12 352 -> 11 554 (ceiling 15 000; 10 886 before the naked-eye stars), cross-catalogue pairs under an arcsecond 23 -> 23 (ceiling 100). The gate's comment now accounts for the survivors by magnitude. gzip -9 sizes against the catalogue before both changes: stars.bin 4 711 922 -> 5 030 741 B, stars-meta.bin 2 576 213 -> 2 784 614 B, stars-index.json 3 724 855 -> 4 003 584 B (+806 KB, 7.3 %). Boot on the dev server, five interleaved cold runs: the task that indexes the catalogue after the data lands, median 1 072 -> 1 182 ms; HUD shown, median 2 622 -> 2 716 ms. This machine measured 0.82-1.30 s for the same baseline task today, above audit #25's 627-843 ms. The drawn-star budget is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
c64eea0803 |
Keep every star the naked eye sees, Rigel, Deneb and Alnilam among them
The 250 pc cutoff took 1 543 of HYG's 8 920 stars of V 6.5 or brighter:
1 339 that both surveys put past it and 204 HYG has no distance for. Five
of the fifty brightest stars in the sky were gone (Rigel, Deneb, Alnilam,
gamma-2 Vel, Wezen), and Naos, Sadr, Aludra and Arneb with them, while
11th-magnitude Gaia stars at the same distance were drawn. The cutoff
bounds a download, not what the sky shows.
placementDistancePc now keeps a star of V 6.5 or brighter at any distance,
at the better of its two distances as before: Gaia's where the archive's
Hipparcos cross-match has a usable parallax, Hipparcos's otherwise. Gaia
saturates on the brightest, so Rigel (264.6 pc), Deneb (432.9), Alnilam
(606.1), Wezen (492.6), Naos, Sadr, Aludra and Arneb sit at their
Hipparcos distance. 1 502 HYG rows come back, 165 of the 206 without a
Hipparcos distance among them because Gaia measured them; 36 fold into a
Gaia entry. 41 naked-eye stars stay out because neither survey gives them
a distance: beta Phe, Polis, Mu Cep, Rho Cas, Eta Car, Alp Cam, Phi Cas,
Chi Aur, Psi-1 Aur, theta-1 Ori, Omi-1 Cen, 66 Ori, 16 Sgr, 10 Sge and 27
HD stars.
Measured on the rebuilt catalogue: 425 117 stars (+1 466), 8 301 of them
past 250 pc (+1 466), the same 336 within 10 pc and 3 652 within 25 pc.
Merge gate: 12 352 HYG rows without a Gaia counterpart (was 10 886; the
ceiling of 15 000 is unchanged, and its comment now counts the naked-eye
stars among the survivors) and the same 23 cross-catalogue pairs under an
arcsecond. Five more exoplanets find their host: HD 81817 b and c,
HD 158996 b, HD 208527 b, HD 220074 b. gzip -9 sizes: stars.bin
4 711 922 -> 4 728 315 B, stars-meta.bin 2 576 213 -> 2 587 977 B,
stars-index.json 3 724 855 -> 3 731 763 B.
The brief behind this also asked to cut once on the best distance, which
would drop the 6 835 stars Hipparcos puts inside 250 pc and Gaia outside.
Those already sit at Gaia's distance (
|
||
|
|
b7f277ea04 |
Answer the review: a short answer makes more survivors, and must not be skipped
Three things this got wrong. The direction: truncating Gaia leaves the HYG rows whose counterpart it dropped without one, so survivors rise — 10 886 today, 12 711 at half the rows, 16 258 at a third — which the comment claimed was the other way, and which decides whether the 15 000 ceiling can be leaned on at all (it catches a truncation past about two thirds, and nothing shallower). The throw: `fetchStars` catches everything a source throws and skips it, so a truncated CSV was reported as "the archive was unreachable" one step after `writeStarAssets` had already overwritten the published catalogue. Marked with `GaiaAnswerError` and rethrown there, so an answer that cannot be worked with fails the run where it happened. Measured end to end in a throwaway working directory, 300 000 rows in the cache: fails, names the cache file to delete, assets untouched. With the rethrow taken back out again: assets written, then "the archive was unreachable". The row limit: `rows.length >= ROW_LIMIT` is true for every reduced ETL_GAIA_ROW_LIMIT, so the tripwire fired on exactly the deliberate slice the override exists for — and told the operator to raise it. Gated on the same flag as its neighbour. `ETL_GAIA_ROW_LIMIT=20000` now runs through; without the gate it dies on the limit it was given. Also: the row floor names the one cache file it is about rather than a glob that takes the Hipparcos cross-match with it, and says an edited query is a third reason it can fire — DEFAULT_QUERY_ROWS now sits under the query it counts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
44f6a8d086 |
Refuse a Gaia answer that came back short, and read the body inside the retry
The merge gate asks whether Gaia contributed any stars, never how many. The TAP service truncates on its own timeout and still serves a well-formed CSV with a 200, ordered by magnitude — so a half answer is the bright half, which is the half HYG overlaps. Every gate passes: Gaia stars are present, HYG survivors go down rather than up, unmerged twins can only fall. The weekly job would publish a catalogue missing two hundred thousand stars and the runner would cache it for the weeks after. `fetchGaiaStars` now refuses fewer than 95% of the 412 765 rows its query holds, as its sibling query already did, and refuses an answer that fills the row limit. `fetchText` retried the request but not the body: a connection reset part-way through the 57 MB CSV rejected out of the loop, with no wait and no second attempt. The read now happens inside it. Also corrected: the merge gate's account of the HYG survivors (two thirds of them are stars Gaia measures but the main query never downloads, since Gaia puts them past the 250 pc cutoff), and the refresh workflow's comment on what happens when the archive is unreachable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
29d3ddb6ef |
Say so when Gaia is missing, rather than as 68 000 unmatched stars
With the merge gate in place, a Gaia DR3 outage no longer ships a HYG-only catalogue: the ETL skips the unreachable source, and validateMerge then fails on the survivor count. That is the right outcome and the wrong message: "68 000 HYG stars found no Gaia counterpart" sends the reader looking at the merge. Gaia contributing nothing is now checked first, by name. Two comments said Gaia was best-effort, in data-refresh.yml and on the merge in fetchStars. They now say what happens instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
f935bae3b3 |
Fail the ETL on a merge that keeps the same star twice
The catalogues are regenerated by a scheduled job that pushes straight to main once the unit suite and a production build pass in the same run. Both passed, every Monday, on a catalogue that carried 23 000 stars twice: the suite tests code against fixtures, and no fixture is 400 000 real stars. Nothing between the ETL and the map ever looked at what came out. Two numbers now have to hold, and each is the signature of a way the merge has actually failed here. Different catalogues placing a star within an arcsecond of each other is never two stars at this depth, and one catalogue does not list a star twice, so every cross-source pair that close is a miss. Nineteen survive today — each a second HYG row wanting a Gaia entry that already absorbed one, which is how Gliese lists some doubles — against 1 112 in the catalogue on main, where a Hipparcos parallax off by half outvoted a direction that agreed to a hundredth of an arcsecond. The ceiling is 100. The epoch failure leaves no close pair at all, because sixteen years of proper motion had already carried the two entries tens of arcseconds apart. What it leaves instead is HYG rows that found no counterpart: 36 056 on main against the 10 876 Gaia genuinely lacks — the stars it saturates on and the red dwarfs past its magnitude cut. The ceiling is 15 000. The pair sweep sorts by declination and walks a one-arcsecond window, so it costs about 300 ms on 423 641 stars — cheap enough to run on every ETL, which is the point: the gate has to sit where the bot already is, before the push, because a GITHUB_TOKEN push fires no CI of its own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
080bbe16dc |
Match exoplanet hosts on the sky, at both epochs the archive might mean
The host cross-reference matched in 3D, nearest star within half a parsec. That is the wrong space for the same reason the star merge learned it: a direction is measured, a distance is inferred. At 170 pc half a parsec is a ten-arcminute cone, wide enough to hand the planets of stars our catalogue does not carry to whatever bright star floats nearest — HATS-6 b sat on HD 39500, seventy arcseconds away. At 60 pc it is tighter than the routine disagreement between the archive's Hipparcos distances and our Gaia ones, which is how four bright giants (7 CMa, HD 81688, omi UMa, xi Aql) lost their planets and GJ 15 A's landed on a neighbouring entry. Hosts are now resolved like stars are merged: by name first, then the nearest star on the sky within a transverse budget — angle times the archive's distance, 0.01 pc — whose distance does not flatly contradict the archive's (the merge's own 50 % ratio). The budget is transverse because the dominant error is proper motion over an epoch difference, a physical displacement that is the same in parsecs at every distance: as an angle it is 60" for Proxima and 2" for a host at 100 pc. Measured on the 504 hosts whose archive name matches a catalogue name outright, true pairs reach 3.4e-3 pc; shifting every host a quarter of a degree finds nothing else within 0.01 but Proxima's own entry, whose budget at 1.3 pc is wider than the shift. The archive never says which epoch a position is for, and they are mixed: alf Tau and GJ 273 publish J2000, HD 133131 and TOI-2459 publish Gaia's J2016. So the query asks for sy_pmra/sy_pmdec too, tries each position at both ends of those sixteen years, and judges a star on whichever is closer. Guess one epoch and a fast star's planets land on a companion: J2016 puts Aldebaran's on Gl 171.1B, J2000 puts GJ 15 A's on a Gaia entry 15.9" out. 1 972 of 6 354 planets now sit on a host, 1 548 before: 432 gained, 26 on a better star (GJ 15 A to Groombridge 34, GJ 676 A off its companion, HD 19994 to 94 Cet), 8 lost — six false 3D matches to stars the catalogue never contained, and GJ 273 b/c, whose archive row says 5.92 pc for Luyten's Star at 3.79: a distance in flat contradiction is exactly what the ratio guard exists to refuse, and the number to fix is upstream. The 2 pc "rematch" apparatus is gone. build.ts recomputed every match after fetchExoplanets had already written the file — at a different tolerance, so the log reported a match count the data did not contain — and the offline entry point that persisted it had no caller. One matcher, one set of constants, used once. The archive cache is now keyed by a hash of the TAP query, so a response cached before the proper-motion columns cannot serve rows without them, where a missing cell would quietly read as "does not move"; the row's astrometry is stored with each planet, which is what made these tolerances measurable offline in the first place. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jxMkwA2rbicdGxHosecYi |
||
|
|
efa9e4084a |
Draw the whole catalogue, and build the aggregation the rest would need
Two things, one verified and one that cannot be. The render budget is now the whole catalogue: 68388 stars, one instanced draw call, which is what a GPU should be asked to do. The budget itself stays, because the catalogue is meant to grow past what any machine should draw at once — Gaia alone could contribute a million — and at that point the selection is what keeps the field legible rather than a grey wash. A `?stars=` override handles the machines that cannot, including the software rasterizer the end-to-end suite runs against, whose frame rate is two orders of magnitude below a real GPU's and which was measuring the rasterizer rather than the app. The aggregation is the second thing, and none of it has run. Every ESA, NOIRLab, SDSS and Euclid endpoint is unreachable from here — only GitHub raw is, which is why HYG and OpenNGC are the current sources. So this is infrastructure and a Gaia query written against the published DR3 schema, not data. What the framework encodes is that these surveys are not interchangeable. The distinction is not size but whether a catalogue knows how far away its objects are, because a 3D map cannot place a star it only has a direction for. Gaia is the only one of the five that can add stars here, because it is the only one that measures parallaxes. DECaPS2 has fifty times Gaia's object count and photometry alone — not one of its 3.32 billion objects can be placed in depth. Euclid's bulge is 8 kpc away, where a parallax is microarcseconds; its contribution would be imagery. SDSS-V and SAGA are keyed to stars something else already places, so they enrich rather than extend. Those roles are recorded as data the ETL prints, not as prose that can drift. Overlapping catalogues are reconciled on direction rather than on 3D proximity, which is the one non-obvious part. Two surveys agree on a star's direction to within an arcsecond and disagree on its distance by tens of per cent, so a star at 200 pc is 50 pc from itself between catalogues while being unmistakably the same object. Matching in 3D would need a tolerance so loose it swallowed real neighbours. The better parallax wins where both reach; where only one does, the star stays. Names become dense-with-holes with a source dictionary, because a survey catalogue has no proper names — writing "Gaia DR3 4472832130942575872" once per star would cost 25 MB per million to repeat what two adjacent fields already say. An empty entry costs three bytes and is regenerated on load. The Sun needed its own case in the merge: it sits at the origin, has no direction to compare, and appears in every catalogue. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G |
||
|
|
29fd92d118 |
Widen the star catalogue, and separate what is drawn from what is known
The map held 8750 stars within 50 pc and rendered 371 systems. Both were lower than they needed to be, for different reasons. The star catalogue was capped by its own encoding as much as by the cutoff: one JSON object per star, eight key names repeated each time, 157 bytes a star. At the range HYG actually reaches that is 17 MB to download and parse before the first frame. So the numbers move into two binary column stores — positions in stars.bin, which the GPU is handed verbatim, and id/magnitude/colour/spectral index in stars-meta.bin — and the JSON keeps only the strings, with 2600 distinct spectral classifications collapsed to a dictionary. The layout is defined once, in star-catalog.ts, and the ETL and the app both use it, so the writer and the reader cannot drift. The cutoff then goes to 250 pc: 68388 stars, 7.8x as many for 1.7x the bytes. That is where HYG's measurements stop rather than a round number — 98.6% of its rows are Hipparcos, whose parallaxes are good to about a milliarcsecond, so beyond 250 pc it would be plotting noise. Drawing all of them is a separate question from knowing them, and it is answered separately. The field draws a budget: every star inside 25 pc, because the nearest are faint red dwarfs and Proxima Centauri is magnitude 11, then the brightest of everything beyond. Search, navigation and the planet cross-reference still see the whole catalogue. A real GPU would draw all 68388 without noticing; the budget is for the machines that would not, and it is one constant. Systems were limited by something else entirely. The archive data already shipped named 4735 host stars and only 388 resolved, because the rest lay outside a 50 pc catalogue — and the cross-reference kept only its own result, so redoing it meant re-downloading an archive that is not reachable from here. Host coordinates are now stored with each planet, and the match is re-resolved at build time against whatever catalogue the run produced. Even name matching alone, which needs no coordinates and so works on the records already shipped, rescues 335 planets across 238 systems: 371 renderable systems become 609. Two selection rules were tuned for a 50 pc bubble and no longer fit. Tethers followed the Sun's nearest neighbours, which are a speck at this range, and now follow the brightest; labels were ranked by proximity, which named whatever sat nearest the middle of the screen, and are now ranked by brightness — so the view names Canopus, Achernar and Spica rather than a clump of catalogue designations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G |
||
|
|
8d8c65bdb2 |
Propagate exoplanets with their real orbital period
Every exoplanet was propagated with gmForParent(undefined) — the Sun's gravitational parameter — so the whole catalogue orbited as though each host were exactly one solar mass. Most hosts are red dwarfs far lighter than that, and a heavier central mass pulls harder and shortens the period, so their planets were whirling round much too fast: TRAPPIST-1 is 0.09 solar masses, and its planets were completing an orbit in roughly a third of the true time. pl_orbper was already in the TAP query and was being discarded on the way into the record. It is now kept, along with st_mass. A period and a semi-major axis together pin the host's gravitational parameter exactly, via GM = n^2 a^3 — no stellar model, no assumption, just the inverse of the orbitalPeriodDays helper that was already there. resolveGravitationalParameter picks the best available source: the measured period, else the published host mass, else one solar mass as before. A derived value implying something outside 0.01-150 solar masses is rejected and falls through, since a period and axis taken from disagreeing solutions would otherwise send a planet spinning at a visibly absurd rate. Note the direction of the error, which is the opposite of what it looks like: assuming a *heavier* host than reality makes a planet orbit *faster*. A test pins it, and caught me stating it backwards first. The NASA Exoplanet Archive is unreachable from this environment (egress policy returns 403 on CONNECT), so exoplanets.json cannot be regenerated here and still carries no periods. Behaviour is therefore unchanged until `npm run etl` is run somewhere with archive access, at which point every planet with a published period starts moving correctly with no further code changes. build.ts reports how many records gained a period, and rejects non-positive ones. Tests: 171 passing, up from 151, including a new end-to-end check that TRAPPIST-1 b with its real period completes exactly one orbit in 1.51088 days and sits a full diameter away at half that. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G |
||
|
|
06cf7d2a15 |
Fix four defects a user hits in the first minute
Found by surveying the codebase against the plan; each was verified against the
committed assets or the running app before being touched.
TRAPPIST-1 was orbiting the Sun. The Exoplanet Archive leaves sy_dist blank for
some systems, and fetchExoplanets.ts read it with bare Number() — Number('') is
0, which is finite, so it slipped past the Number.isFinite guard in
resolveHostStarId, placed the host at the origin, and matched Sol at distance
exactly 0. 127 records shipped with hostStarId 0, all seven TRAPPIST-1 planets
among them, and the system view filters on that id, so drilling into Sol drew
127 alien worlds inside the real solar system.
Fixed in three places: resolveHostStarId now rejects a non-positive distance
(the robust guard, covering every caller), fetchExoplanets.ts uses the
parseOptionalNumber that already sat unused in that same file for ra/dec/dist,
and validateExoplanets asserts nothing ever resolves to the Sun again — the Sun
has no exoplanets, so that tripwire costs nothing and is permanent.
The archive's endpoint is blocked by this environment's egress policy, so the
ETL cannot be re-run here. The committed asset was corrected in place instead,
which is safe because the outcome is deterministic: the name path runs first and
none of the 127 resolve by name, so all of them reached id 0 positionally and
the fixed pipeline yields null for exactly that set. Cross-referenced hosts drop
from 761 to 634; record count is unchanged.
Dragging to rotate selected stars. Selection was bound to the raw click event,
which browsers fire on release however far the pointer travelled and which
OrbitControls does not suppress — so any drag ending over a star launched a
camera flight, and in system view routed away to /body/:id. Now tracks
pointerdown and ignores a release more than 5 px from it.
Ghost systems accumulated on every star-to-star hop. SystemOrbitsRenderer.dispose
released geometries and materials but never detached its group, so old orbit
lines stayed parented forever — still traversed and re-uploaded each frame with
disposed geometries, drawn over the new system and unpickable. dispose() now
detaches and clears.
Galaxy star labels stayed pinned inside the system view. They are CSS2D objects
parented to the scene rather than to galaxyGroup, so hiding the group left up to
15 parsec-space names clumped over the system's star. Cleared on entry. Also
gated the per-frame Kepler propagation on actually being in a system; it ran in
galaxy view too, because the renderer is never nulled on exit.
Tests: 116 passing, up from 112. Build, both typechecks and the Playwright suite
are green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G
|
||
|
|
3a859360ba |
Add the deep-sky backdrop, the last unbuilt piece of the plan
The design doc scopes deep-sky objects as a galaxy-view backdrop and lists fetchDeepSky.ts, deepsky.json and deepsky.model.ts, but none of it existed — it was the only part of the plan with no implementation behind it. ETL: fetchDeepSky.ts pulls the OpenNGC catalog, classifies each object as a galaxy/nebula/cluster, and keeps the ~460 worth drawing (everything Messier, everything with a common name, and anything brighter than magnitude 9) out of ~12,000 mostly-anonymous rows. build.ts runs it and validates the output. Distances are the hard part: OpenNGC has no distance column, and both fallbacks fail for the best-known objects. M31, M33 and M42 are Local Group members whose redshift is negative or absent, and a galaxy's catalog parallax comes from a cross-matched foreground star — 6 mas for M31 would put a 780 kpc galaxy at 167 pc. So records store a unit direction on the celestial sphere rather than a position (the line of sight is always known precisely, and the objects are drawn on a fixed backdrop shell where true distance is unusable anyway), and distance is optional metadata carrying its own provenance. Parallax is trusted only for galactic objects, redshift only above z=0.003 where expansion outweighs peculiar velocity. 330 of 463 get a distance; the rest honestly report none. Rendering: DeepSkyRenderer paints the objects as soft additive billboards on a 2500 pc shell — clear of the 50 pc star field, beyond the camera's 2000 pc orbit limit, and inside its 5000 pc far plane. Size comes from real angular extent, so Andromeda is six times wider than the full Moon, clamped at both ends. Sprites rather than points because the WebGPU backend caps point primitives at one pixel; materials are shared per kind and brightness band, so 460 objects cost nine of them. The brightest dozen get permanent labels, which needed the label overlay to accept string ids alongside numeric star ids. The backdrop is decorative, so a failure to load its dataset is logged and the star field comes up regardless. Also documents the app in the README, which until now covered only the plugin marketplace. Tests: 112 passing, up from 54. Build, both typechecks and the Playwright suite are green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G |
||
|
|
d7e8ea1d4d |
@
Add star-map Angular app, ETL pipeline, and caveman plugin Angular 3D star map (galaxy/system/body views, Three.js rendering, navigation store) plus the NASA ETL tooling that builds the star, exoplanet and solar-system datasets, Playwright e2e suite, and the cs:caveman Claude Code plugin (command, agent, skill). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> @ |