Most news is noise. A fire, a typhoon, a celebrity, a court case nobody will remember next week. But buried in the same feeds are the few changes that move money for years — a new licence regime, a levy, a subsidy scheme, a consultation paper. I built a machine to find those every morning before I wake up.
This is how it works, what it costs, and the mistakes that only showed up once it was running for real.
The Idea
I wanted one page every morning that answered three questions:
- What changed? A key regulation, a major business event, or a lawsuit that creates an opportunity.
- How does the money flow? Who pays, who receives, how big the pool is, and when.
- What should I start doing about it? A concrete first step, not "keep an eye on this".
Across Hong Kong, mainland China and the rest of the world, every day.
The Test That Decides Everything
Most news scanners filter by topic: keep "finance", drop "sport". That keeps loud trivia and drops quiet structural change. A gazette notice about port fuel rules has no exciting keywords, but it can matter for years.
So the filter asks a single question about every headline:
"Does this change a rule, a price, a capital flow, or a market structure in a way that still matters in six months?"
// The persistence testA traffic accident: no. A new stamp duty: yes. A CEO resigning: only if it signals a change of strategy. A court ruling: yes, if it sets a precedent. And when the model is unsure, it keeps the item — a missed regulation change costs far more than one extra analysis.
What I Actually Built
A Python pipeline that runs on my PC at 8am, Hong Kong time. Each stage is a separate step, so a crash in one never corrupts the others and a failed run can simply be re-run.
Step 1: Collect from the right places
The machine reads 47 feeds: Hong Kong government press releases, RTHK, Hong Kong Free Press, the US Federal Register, SEC and European Commission releases, major business news, mainland financial news, and about 25 Google News keyword searches like 香港 牌照 制度, 香港 補貼 資助 and 大灣區 政策.
Two early findings shaped the design. Google News links cannot be opened without running a browser, so I use Google News only to measure how many outlets carry a story — a free importance signal. And for regulation, the Hong Kong government's own press releases are better than the news: around 100 official announcements a day, in Chinese, published before any journalist writes them up. The news tells you the reaction. The government feed tells you the event, earlier and complete.
Step 2: Throw away the obvious, for free
Before any AI is involved, simple rules drop stale items, weather, traffic, crime, entertainment and duplicates. On a typical morning that turns 3,076 headlines into 449 without spending a cent.
Step 3: Ask a cheap model the persistence question
The 449 survivors go to a small, cheap model in batch mode, which runs overnight-style at half price. It labels each one: does it last, and what kind of change is it? On October 3rd, 166 passed, and the categories tell you what a morning in Hong Kong actually looks like:
| Category | Items |
|---|---|
| Regulation | 49 |
| Macro policy | 33 |
| Market structure | 23 |
| Subsidy or funding | 13 |
| Infrastructure | 11 |
| Corporate action / trade & sanctions | 10 each |
| Tax or levy | 6 |
Some of the items that passed that morning: statutory minimum pay for foreign domestic workers rising to HK$5,220 a month, a new vaping tax, a 30% rise in minimum alcohol pricing, proposed custody rules for crypto assets, and a gazette amendment to port rules for green ship fuel. None of them made the front page. All of them change a price or a rule for years.
Step 4: Remember yesterday
A story that appears every day for two weeks is not news on day ten. So every surviving item is linked to a "story ledger" — a small database of everything the machine has seen, tagged as new, developing, accelerating, quiet or dormant. On October 3rd the 166 items grouped into 124 stories, and only 39 of those were genuinely new. Only new or materially updated stories go further.
Step 5: Score against my own situation
"What should I start doing" is useless if the model doesn't know who I am. A private profile file describes the sectors I know, what I can actually execute, and what I never want recommended. The model scores each story for fit, and things like "African heads of state break ground on a $16 billion refinery" score close to zero — huge, but nothing I could act on.
Step 6: The deep analysis
The top three stories go to a larger model, which has to return a strict structure for each:
- Classification — type, size, time horizon, why it lasts, and what would prove it wrong.
- Money flow map — who pays, who receives, how big the pool is, when — plus the second-order flows everyone misses, and who loses.
- Actions — each tagged as capital, operating, venture or watch, with a time window, a cost, and a concrete first step.
Then it renders a page in this site's neon style, in Traditional Chinese, and publishes it here.
The Prompt That Wrote Finance Essays
The first version of the money-flow prompt produced beautiful paragraphs that said nothing. "This regulation may create opportunities for compliance providers and could affect market participants." True of every regulation ever written.
It took several rounds of tightening to force it to name actual payers, actual recipients and actual amounts — and to say plainly when the source did not give a number instead of inventing one. The difference looks roughly like this:
The Bugs That Only Real Mornings Found
// The feed with no dates
One mainland financial feed published 288 items with no publish dates at all. My "ignore anything older than a few days" rule could not see them, so months-old articles looked fresh and quietly became 37% of everything the machine kept.
// The abandoned mirror
Another feed came through a mirror site that had stopped updating. Its newest item was 256 days old. It had passed every check because I only checked that it returned items, not that they were recent.
// The scheduled task that couldn't read Chinese
Every test run worked. The first unattended run crashed on the first Chinese source name, because Windows' scheduled tasks use a different text encoding from my terminal. It now forces UTF-8 itself instead of trusting how it was started.
// The deploy that silently didn't
The website host only deploys commits it can attribute to the account owner. One wrong email in an automated commit and the page simply never goes live, with no error on my side. So the identity is set explicitly on every commit and checked after each push.
// The uncomfortable bit
Only the encoding crash failed loudly, and that was the easy one to fix. The others still produced reports that looked completely normal. The undated feed made the brief worse, not broken; the stale mirror added plausible-looking old news. An automated system's most dangerous failures are the ones that still produce output. Now any new feed has to pass two checks before it is added: enough items, and items that are actually recent. Undated feeds are rejected outright.
Why the Reports Stay Unlisted
The reports live on this site at an unlisted address that asks search engines not to index it. That was a deliberate choice, made before the first page went live:
- Your edge becomes public. A daily opportunity brief is read by anyone, including people who can act faster than you.
- News feeds have terms. Google News RSS is meant for personal reading. Original analysis is fine; republishing headlines at scale is a different thing.
- Investment advice is regulated. In Hong Kong, advising on securities needs an SFC licence. My notes to myself are not the same as recommendations to the public.
What It Costs
I originally planned it around Claude, then built it on OpenAI's API because that was the key I had. The split stayed the same: buy judgement cheaply for the bulk filtering, and pay properly for the few analyses where all the value is.
| Stage | Volume | Model |
|---|---|---|
| Free rules | 3,076 → 449 items | None |
| Persistence test | ~450 items, in batch at half price | Small model |
| Relevance scoring | ~40 new stories | Small model |
| Deep analysis | 3 stories | Larger model |
| Total per day | ≈ US$0.37 | |
About eleven dollars a month to have someone read 3,000 headlines before breakfast.
What I Learned
Filter by consequence, not by topic
The persistence question — does this change a rule, a price, a capital flow or a market structure for six months? — does more work than any keyword list. It keeps the boring gazette notice and drops the dramatic headline, which is exactly backwards from how news is presented.
Go to the source, use the news for salience
Government feeds give you the event first. News coverage gives you a free importance signal: how many outlets picked it up. Each is good at a different job.
Make vague answers impossible
Language models default to sounding helpful. A strict output structure that demands names, amounts and dates — and accepts "not disclosed" — turns essays into something you can act on.
Check freshness, not just existence
A feed that returns 288 items is not healthy if none of them have dates. "It returned data" is the weakest possible test of a data source.
Spend where the value is
Ninety percent of the reading is cheap filtering. Ten percent is real analysis. Paying flagship prices for the first part, or bargain prices for the second, are both mistakes.
Was It Worth It?
It has run every morning since the end of August — more than thirty reports so far. Most days, two of the three stories are interesting and one is a stretch. But the habit it built is the real value: I now start the day with three changes that matter, the money behind them, and one thing to do, instead of scrolling through the noise to find them myself.
It is a research tool for my own decisions, not investment advice. And like everything I build lately, the hardest part was not the AI. It was deciding precisely what question the AI should answer.
Back to Blog