The machine at a glance
One clock, five actors, one direction of flow, and a feedback loop that runs the other way. The full inventory of each stage follows below.
The whole machine on one clock, read from the repository at each build. Every routine here is a step-by-step chart on the algorithm page, where a click flies into it and each step shows its criteria, what it reads and what it writes.
The 24-hour clock (UTC)
The machine runs on a fixed daily rhythm. Collection lands before thinking; thinking lands before publishing.
| 02:20 | housekeeping | Stale-branch cleanup | closes abandoned work branches |
| ~03:30 | thinking | Daily brief | weekdays — sweeps every lane, writes the brief, picks the lead items |
| 05:30 | collection | Earnings + market data | Finnhub calendar and prices; Mondays also the weekly IR watch |
| 06:00 | collection | Orbit track | GRAVITAS altitude history from TLEs |
| 08:00 | publishing | Site deploy | also redeploys on every merge to main |
| hourly :17 | feedback | Telegram feedback capture | reader verdicts land in feedback.md |
| 20:30 | collection | Source mirror, run 1 | all sources, ahead of the evening routines |
| 20:50 | collection | Tripwire hit log | commits the day's live alerts |
| 21:00 | thinking | Daily retro | weekdays — scores the brief against reader verdicts, rewrites doctrine |
| ~22:00 | thinking | Dossier deep-dive | daily — moves one dossier forward |
| 23:00 | collection | Source mirror, run 2 | the snapshot the next morning's brief reads |
| Sun 09:00 | housekeeping | Ops repair | weekly — fixes whatever the week's retros flagged |
| Sun 10:00 | extraction | Transcript tone scoring | FinBERT over the week's earnings calls |
The full inventory, stage by stage
1 Collect — 67 mirrored sources, plus the live layer
Twice a day (20:30 and 23:00 UTC) the source mirror fetches every watched source into sources/latest/ with a manifest of hashes, so each morning's brief reads a known snapshot. A separate live layer never sleeps.
The mirror, by family
- SEC EDGAR — filing feeds for 14 tracked issuers (SES peers, operators, suppliers), plus full-text search for satellite 8-Ks.
- FCC — the Space Bureau's daily documents (the source of the SAT- filing lane), five ECFS docket watches (C-band, market-access reciprocity, MSS spectrum), and the headlines feed.
- ITU — the as-received space network register, every filing's detail page (found by scanning publication ids directly, because the register only serves its first page), the bringing-into-use register, and the SNS technical databases each filing links.
- Trade and specialist press — about twenty feeds: SpaceNews, Payload, Space Intel Report, Breaking Defense, china-in-space, European Spaceflight and peers. One (Via Satellite) blocks fetches; a standing search pass compensates.
- The X mirror — eight analyst and journalist handles read for genuinely new claims, then chased to primary sources before use.
- Investor relations — SES, Eutelsat and Viasat results and regulated-information pages; the Finnhub earnings calendar; daily prices.
- Launch logs — planet4589 and china-in-space, the ground truth for constellation deployment counts.
- Policy — RSPG consultations and EC defence-space actions.
The live layer
- Tripwires — a Cloudflare Worker watches 14 fast-moving feeds (EDGAR in real time, the ITU register, the FCC Space Bureau, trusted breaking-news handles) and alerts the moment one moves. Hits are committed to the repository nightly.
- Reader inbox — transcripts, decks and documents dropped by hand; PDFs convert to Markdown automatically on arrival.
- Telegram — every reader reply is captured hourly into
feedback.md. A reply is a verdict, and verdicts steer the machine (see the loops). - Freshness rule — if the snapshot is more than four hours old when a routine starts, live refetching becomes mandatory, not optional.
2 Extract — parsers that refuse to guess
Extractors turn mirrored pages into structured records. Every one ships a self-test with negative controls that are proven to bite, and every merge is additive — an extractor may add or update rows, never silently drop them.
The spectrum lane (the most built-out)
extract-sat-filings.py— FCC Space Bureau notices → who filed, what kind, which bands; also detects processing-round public notices and their cut-off dates.extract-itu-filings.py— the ITU register and detail pages → one row per filing, with objections ("Spacecom Comments") separated into their own file, because a comment is not a filing.extract-itu-biu.py— joins bringing-into-use records onto filings, using only the regulator's own stage vocabulary.extract-itu-mdb.py— downloads each filing's SNS technical database, reads orbital shells, frequencies, channel widths and gateway sites, and deletes the binary. Nothing binary is ever committed.build-systems.py— the assembler: every lane above collapses into one canonical corpus,systems.jsonl.
The other lanes
- Company desk — EDGAR indexes, XBRL company facts and transcript indexes per tracked company.
- Series — launch cadence and constellation-deployment counts, from the launch logs.
- Tone — FinBERT scores every earnings-call transcript weekly for tone, stance and analyst reaction.
- Orbit — one satellite's altitude history, tracked daily from public TLEs.
- Lint suite — briefs, dossiers, citations and repository layout are machine-checked on every change, and every rendered sentence passes the editor gate before it can merge.
3 Canonical stores — one place per fact
The design rule: a fact is derived once, stored once, and every consumer slices the same store. When a value looks wrong on a page, the bug is upstream — never patched in the page.
- The filings corpus —
data/spectrum/systems.jsonl: 687 satellite systems assembled from 1048 ITU filings, 219 FCC filing records and the technical databases. The Filed board, the daily brief and Quick-Bytes all read these same rows. 2318 ITU objections live separately. - The company desk —
data/companies/: filings, financials and transcripts for 31 companies. - Series —
data/series/: launch, constellation, funding, financials and correlation series behind the charts. - The house view —
priors.md: the threads, the standing entity watchlist, and the doctrine rules. Rewritten only by the evening self-review. - Dossiers — 18 thematic dossiers and 28 company profiles, each a claims table where every row carries a date and a source.
- The record — 93 daily briefs, the scorecard, the feedback log, the events calendar and the dockets watch.
Honest status: the one-store pattern is fully enforced on the filings lane, and the financials corpus (data/financials/financials.jsonl) now collapses the filing-derived and desk-compiled metrics into one store. Pages still read the older stores directly — migrating them onto the corpus is the next architectural move.
4 Think — seven scheduled reasoning routines
The stores do not interpret themselves. Seven scheduled routines — each a written procedure in routines/, executed by a reasoning model — do the thinking, on the clock shown above.
- Daily brief (weekdays, before the European morning) — sweeps every lane plus live fetches, a Mandarin-language pass for Chinese programs, and the X mirror; writes the brief; picks the lead items that become this site's front page. Weighs every item through the house view.
- Daily retro (weekday evenings) — scores the morning's brief against reader verdicts and reality, logs every miss with a root cause, and rewrites
priors.md. The doctrine page shows its revision trail. - Dossier deep-dive (nightly) — advances one dossier: verifies claims, compresses over-budget sections, hunts what the day's coverage missed.
- Ops repair (Sunday) — fixes the pipeline itself: dead feeds, broken parsers, whatever the week's retros flagged.
- IR watch (Monday) — the investor-relations pass over the week's results and filings.
- Deep-dive (on demand, gated) — a long-form investigation when a question earns one.
- Quick-Byte (on demand) — one event, distilled against everything above into a one-screen executive email, built to a written archetype per event class.
5 Publish — where the work surfaces
- This site — rebuilt on every change and daily at 08:00 UTC. The front page is the latest brief's lead items; the board and thread pages render the house view; the fleet page's Filed board renders the filings corpus (the spectrum page tracks dockets); fleet, companies, dossiers, signals, IR, calendar and search each slice their store.
- The daily brief — delivered to Telegram, where replies become the feedback that steers the machine.
- Quick-Bytes — executive email, sent through the same pipeline that archives every byte in the repository.
- Tripwire alerts — live pushes the moment a watched feed moves, hours ahead of the daily cycle.
The editor gate
Between thinking and publishing sits an editor that never tires. Every routine writes to one prose contract: lead with who did what, one idea per sentence, citations in their own field, nothing about the desk's own process. 3 scripts check it, using 24 written rules. They run inside the routine before it opens a pull request, then again in the merge action. One error blocks the merge and sends one message to the phone. Nothing reaches this site until it passes.
briefs/lead items, weak signalspriors.mdthread blurbscloud-synopsis.mdthreads-page narrativedossiers/dossiers, company profilesevents.mddated watch-pointsdockets.mdlive proceedingsir-ledger.mdmarket overhangsThe rules every lane shares
- no sentence over 40 words
- no semicolon joining two claims
- two asides per sentence at most
- no numbered lists inside a sentence
- citations are links, never brackets
- no "this desk", "this cycle", no file names
- every fact, date and hedge survives the edit
Not gated, on purpose: the numbers. Prices, series, the filings corpus and the earnings Q&A are data or verbatim record, not prose. The wires are external headlines as captured. The scan queue is a note to the operator, not to the reader.
The loops inside the loops
The big loop is the daily cycle. Inside it, four smaller loops make the machine self-correcting.
A reader replies to any brief item on Telegram → the reply is captured hourly → that evening's retro scores the item against the verdict → the score changes priors.md → the next morning's brief weighs differently. A one-line reply reshapes tomorrow's judgement.
Every miss the retro finds gets a root cause. A doctrine miss changes the rules the same evening. A pipeline miss (a dead feed, a parser gap) goes to the ops backlog and Sunday's repair pass fixes the machine itself.
Claims carry sources; a verify-before-citing list tracks every claim that once went wrong and how; dossier rows are corrected in place with the correction dated. The record shows its own repairs.
Every extractor runs its self-test before touching data; merges never wipe; a changed page is verified by its content, not its status code; and absence of evidence is recorded as absence, never as fact.
Rules the whole machine obeys
- Primary sources outrank coverage. A press story is a pointer; the filing, the transcript or the docket is the fact.
- Regulators' own words. Filing stages use the ITU's and FCC's own vocabulary — the machine never invents a status.
- State what a thing is before what it means. Analysis is welcome; it is marked as the house read and never overwrites the record.
- One store per fact. Consumers slice; they never re-derive.
- Show the work. Every claim carries a date and a source; the doctrine page shows every change of mind.
- Readable by contract, not by mood. One prose rule set, checked by scripts, blocks the merge. Density is not detail: a fact in its own sentence is detail.