Infinite Horizons · Tech Stack · Ray Ferguson · Singapore · 29 September 2026
Half the tokens, 14% of the spend
The most useful chart of the month comes from Vercel’s September production index, picked up this week by the Financial Times. Open-weight models are those whose trained parameters are published, so a company can download the model and run it on its own hardware. That is not the same as open source. The weights are yours to use; the training data and the recipe stay with the lab. They now handle more than half the tokens passing through the routing platforms that publish their numbers, for about 14% of the spending.
Those platforms are a slice of the market and a token is a rough proxy for work done, so read it as direction of travel. The direction is clear enough.
AT&T runs about 40% of its AI workloads on open models and intends to reach 70% within a year, at 45 billion tokens a day. On Vercel’s gateway, open weights carried 56% of tokens in August against 7% last December. The frontier labs are cutting prices and still take most of the revenue.
Six months apart, on the same stack
Singapore’s Minister for Foreign Affairs, Vivian Balakrishnan, explained why no government is about to fix this, at the Asia Society last week.
His first point is structural. Where the Sputnik contest ran between two separate technological ecosystems, America, China and everyone else now work on the same mathematics, the same techniques and largely the same machines, and because that stack has been open and competitive, acceleration has been dramatic. His second point is the league table: on AI frontier models he puts America ahead by “maybe six months”, on biotechnology neck and neck, on renewable energy China well ahead.
Six months governs everything else here.
It is too narrow for any government to impose restraint on its own labs, since whoever gets there first takes an economic lead and the hop to military advantage is short. It is also too narrow for boards and senior teams to treat the frontier as a permanent quality premium.
Balakrishnan frames the moment as a new Gilded Age, and notes that the last one was answered with the welfare state, socialism, communism and the extreme right. He asks whether the plot is being recycled.
Trade secret or patent
His most useful point for anyone buying AI concerns open weights. For fifteen years machine learning advanced on a norm of publishing: shared papers, available models, weights others could retrain, and competition through academic recognition. China’s labs continue that norm. The closed American frontier models are the departure from it.
His analogy is the industrial revolution’s answer to the same problem. A closed model is a trade secret, the Coca-Cola recipe kept in the safe. Open publication is the patent bargain: disclose for the common advance, take a temporary monopoly, and let it expire before it stifles anyone else.
We settled that once, for chemicals and machines.
We have not settled it for weights, and that unsettled question is where boards now operate.
The rules are still being drafted
Nobody is about to hand anyone the answer.
Europe’s response to the pressure was to postpone: the Digital Omnibus on AI pushed high-risk obligations from August this year to December 2027, and product-embedded systems to 2028. Across the water Washington is divided, federal preemption of state law is stalled in committee, and California’s governor has until tomorrow to sign or veto the state’s Frontier AI Safety Act. In Beijing the state censors chatbot output while racing ahead on capability.
I have sat in enough rooms to recognise the condition. It is what a senior team reaches when it agrees something matters, agrees it is moving faster than the committee calendar, and agrees to return to it next quarter.
Which models run inside the company, on whose infrastructure, with whose data, is a live question today and will be settled in law later. Where the supervisor has yet to form a view, boards carry the judgement.
Who made the weights
The open-weight market is dominated by Chinese labs. Qwen, DeepSeek, Kimi and GLM all publish downloadable weights, and Chinese models fill most of the published top ten by tokens processed. Mistral and several American labs release on similar terms, but the volume is Chinese.
The order within the pack changes constantly. On the Artificial Analysis index this month, GLM and Kimi sit joint first among open-weight models, with Qwen next. Nine points separate the open-weight leaders from GPT-6 and Claude Fable. For document summaries, code assistance, call transcription and internal search, that distance costs little and the saving is roughly tenfold. A shortlist drawn up in June was stale by September, so what a firm needs is a written policy that survives the next reshuffle.
The question now: what rule governs which weights go where.
What Singapore decided
Singapore answered that in practice while Brussels was still deferring. The posture is soft-law and sector-specific, with binding rules where the stakes justify them, which in my world means MAS. Its Model AI Governance Framework for Agentic AI rests on one principle: humans are ultimately accountable.
Provenance has been settled by conduct. AI Singapore’s SEA-LION v4.5 is built on a Qwen architecture and post-trained for eight regional languages, and NCS, Singtel’s technology arm and the main systems integrator for the Singapore government, has folded Qwen into its products, telling a Singapore audience in July that even government customers are open about it.
Stated as a working rule, and this is my own formulation with no official standing: the weights can be Chinese, the servers must be yours, the data stays inside your perimeter, and the accountable person is named and human. PDPA obligations apply to all of it already, whatever the AI-specific guidance eventually says.
Ownership changes responsibility. AI Singapore’s model card is blunt that SEA-LION has not been aligned for safety or hardened against adversarial prompting. That work falls to whoever deploys it.
The correspondent bank rule
I ran banks in markets where every dollar cleared through New York, and the first thing a decent treasurer learns is that you never run on a single correspondent. You keep more than one, you know which relationships you would lose in a designation event, and you have rehearsed the day it happens. We called it plumbing, and it kept us alive more than once.
Model routing is the same class of plumbing. A company whose inference all runs through one provider has a single correspondent, and so does one whose open weights all come from a single jurisdiction, with a different failure mode. The test of a reserve is whether it survives the failure you are hedging. A second model on the same cloud, behind the same orchestration layer, gives you one correspondent wearing two hats. In July this series followed a model switched off and back on at a supplier’s discretion, and watched Brussels and Beijing turn that episode into policy.
Three lines belong in the minutes. A routing policy, setting which class of data may go to which class of model. A register of every model in production, with its origin, licence, version and the jurisdiction of the lab that trained it. Name the owner. And a tested failover, because a plan that has never been run is a hope. Test the switch.
The honest limit
Singapore’s frameworks are voluntary outwith the regulated sectors, and a voluntary framework holds as long as firms choose to hold it. Cheaper tokens rarely mean a smaller bill either. Open weights change who you pay and how much control you keep, and they leave consumption where it was.
The settlement will come, as trademarks and patents came. Until it does, six months is all that separates the leaders, nobody is slowing down for anyone, and the board is the de facto regulator of every model it runs.
Ray Ferguson writes the Infinite Horizons Substack on leadership and judgement in a messy world. Signal Stack reads geopolitics, markets and structural change. Tech Stack reads AI, technology and cyber. Each edition is written plainly, for board-level readers, with evidence ahead of opinion and no patience for jargon. Full archive of previous Signal Stacks. The wider body of work lives at ferguson.sg. Ray is a chairman and international banker with 35 years of leadership across Asia, the Middle East, Europe and the Americas. Former CEO of Standard Chartered Americas, CEO of Standard Chartered Singapore and SEA, UAE and Oman, and Deputy CEO of Bank ABC, Bahrain. Chair of insurer Singapore Life, 36ZERO Marine Group, and ESGpedia, Asia’s leading ESG data and technology company. Views expressed are personal. Singapore citizen, ocean sailor, Criticaleye board mentor, and author of Infinite Horizons: Seasons in The Sun, Life & Leadership Lessons, publishing with World Scientific in November 2026. It sets out what that career taught him about counsel, stewardship and endurance, and about the discipline required to sustain a life alongside the work. It takes the form of a memoir at its core. All author proceeds go to the Wheelchair Rugby Association (Singapore).


Great post