Widen the star catalogue, and separate what is drawn from what is known

The map held 8750 stars within 50 pc and rendered 371 systems. Both were
lower than they needed to be, for different reasons.

The star catalogue was capped by its own encoding as much as by the
cutoff: one JSON object per star, eight key names repeated each time, 157
bytes a star. At the range HYG actually reaches that is 17 MB to download
and parse before the first frame. So the numbers move into two binary
column stores — positions in stars.bin, which the GPU is handed verbatim,
and id/magnitude/colour/spectral index in stars-meta.bin — and the JSON
keeps only the strings, with 2600 distinct spectral classifications
collapsed to a dictionary. The layout is defined once, in star-catalog.ts,
and the ETL and the app both use it, so the writer and the reader cannot
drift.

The cutoff then goes to 250 pc: 68388 stars, 7.8x as many for 1.7x the
bytes. That is where HYG's measurements stop rather than a round number —
98.6% of its rows are Hipparcos, whose parallaxes are good to about a
milliarcsecond, so beyond 250 pc it would be plotting noise.

Drawing all of them is a separate question from knowing them, and it is
answered separately. The field draws a budget: every star inside 25 pc,
because the nearest are faint red dwarfs and Proxima Centauri is magnitude
11, then the brightest of everything beyond. Search, navigation and the
planet cross-reference still see the whole catalogue. A real GPU would
draw all 68388 without noticing; the budget is for the machines that would
not, and it is one constant.

Systems were limited by something else entirely. The archive data already
shipped named 4735 host stars and only 388 resolved, because the rest lay
outside a 50 pc catalogue — and the cross-reference kept only its own
result, so redoing it meant re-downloading an archive that is not
reachable from here. Host coordinates are now stored with each planet, and
the match is re-resolved at build time against whatever catalogue the run
produced. Even name matching alone, which needs no coordinates and so
works on the records already shipped, rescues 335 planets across 238
systems: 371 renderable systems become 609.

Two selection rules were tuned for a 50 pc bubble and no longer fit.
Tethers followed the Sun's nearest neighbours, which are a speck at this
range, and now follow the brightest; labels were ranked by proximity,
which named whatever sat nearest the middle of the screen, and are now
ranked by brightness — so the view names Canopus, Achernar and Spica
rather than a clump of catalogue designations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WaySiNst4HhDXBHnMy8p5G
This commit is contained in:
Claude
2026-08-05 08:35:41 +00:00
parent 029162ff52
commit 29fd92d118
20 changed files with 692 additions and 49 deletions
@@ -1,4 +1,5 @@
import { CartesianCoordinates, distanceBetween, raDegDecDistanceToXyz } from './coordinates';
import { ExoplanetRecord } from '../models/exoplanet.model';
import { StarRecord } from '../models/star.model';
/** Normalizes a star name for comparison: lowercase, alphanumeric characters only. */
@@ -65,3 +66,68 @@ function findNearestStarWithin(position: CartesianCoordinates, stars: readonly S
return closest ? closest.id : null;
}
/**
* Re-resolves every exoplanet's host star against a star catalogue.
*
* The cross-reference is a *derived* fact: it depends as much on which stars were loaded as on
* the archive itself. When the catalogue reached 50 pc, 388 of the archive's 4735 named hosts
* found a match and the other 4347 were carried and never drawn — not because their planets are
* unknown, but because their star was out of range. Widening the catalogue rescues some of them,
* and until the host coordinates were stored alongside each planet that meant re-downloading an
* archive which is not always reachable.
*
* Records written before those coordinates were kept can still be matched *by name*, which needs
* no coordinates at all — and that alone is worth doing, because a wider catalogue contains more
* names. What such a record cannot do is disprove its existing match: a name miss means only
* that the name missed, not that the star is absent. So those are upgraded where a match is
* found and left alone otherwise, while records that do carry coordinates take the new result
* outright, match or no match.
*/
/** A host must sit within this many parsecs of a catalogue star to count as the same object. */
export const HOST_MATCH_TOLERANCE_PC = 2;
export interface RematchSummary {
total: number;
/** Records carrying host coordinates, and therefore eligible to be re-matched in full. */
resolvable: number;
matched: number;
gained: number;
lost: number;
}
export function rematchHostStars(exoplanets: ExoplanetRecord[], stars: readonly StarRecord[]): RematchSummary {
const nameIndex = buildStarNameIndex(stars);
const summary: RematchSummary = { total: exoplanets.length, resolvable: 0, matched: 0, gained: 0, lost: 0 };
for (const exoplanet of exoplanets) {
const { hostRaDeg, hostDecDeg, hostDistancePc } = exoplanet;
const positioned = hostRaDeg !== undefined && hostDecDeg !== undefined && hostDistancePc !== undefined;
if (positioned) {
summary.resolvable++;
}
const previous = exoplanet.hostStarId;
// With no coordinates the query still carries the host's name, and `resolveHostStarId` tries
// that first; the positional fallback simply declines to run on non-finite coordinates.
const resolved = resolveHostStarId(
{ hostname: exoplanet.hostStarName, raDeg: hostRaDeg ?? Number.NaN, decDeg: hostDecDeg ?? Number.NaN, distancePc: hostDistancePc ?? Number.NaN },
stars,
HOST_MATCH_TOLERANCE_PC,
nameIndex
);
exoplanet.hostStarId = positioned ? resolved : (resolved ?? previous);
if (exoplanet.hostStarId !== null) {
summary.matched++;
}
if (previous === null && exoplanet.hostStarId !== null) {
summary.gained++;
} else if (previous !== null && exoplanet.hostStarId === null) {
summary.lost++;
}
}
return summary;
}