Build to Keep
FIELD NOTES

Rent breadth. Own the decision.

Thinking Machines' new open-weight model clarifies a choice most businesses have not yet made: when to rent general intelligence, and when to control a decision system of their own.

Most model launches invite the same question: how capable is it?

Thinking Machines' release of Inkling invites a more useful one: what, exactly, should a business own?

Inkling is a broad, multimodal model with openly available weights under an Apache 2.0 licence. The company is clear that it is not the strongest general-purpose model available. Its case for Inkling is different. It is a capable base model that organisations can fine-tune for their own use. Source: thinkingmachines.ai

That is not a reason to cancel a frontier-model subscription. It is a reason to stop treating every AI use case as the same purchase.

Two different purchases Left panel: breadth, rented. Lines fan out from one point to many endpoints, a different question each time. Right panel: the decision, considered for control. One line returns to the same mark again and again. BREADTH RENT IT a different question each time THE DECISION CONSIDER CONTROLLING IT the same judgment, thousands of times
Two purchases. Open-ended work rents general intelligence; the repeated judgment is the one to consider controlling.

The first purchase is breadth. Rent it.

Use a frontier model when the work is open-ended, the question changes each time, and the value lies in wide knowledge or general reasoning. Research, writing, coding, planning, and complex analysis all belong here. These models improve quickly, and no sensible business should try to reproduce that rate of progress for itself.

The second purchase is the decision. Consider controlling it.

Businesses do not only ask novel questions. They make the same judgments thousands of times: whether a sales lead is worth pursuing, where a claim should go, how a customer request should be routed, or whether a document meets a standard.

In work like this, general intelligence may be more than you need. What matters is whether the system applies your organisation's judgment reliably.

A model trained and tested on high-quality examples from your business can outperform a larger general-purpose model used with a prompt alone on a bounded task such as classification. That does not mean smaller models are inherently better. It means task fit can matter more than general capability when the job is narrow, repeated, and measurable. Source: arxiv.org/abs/2406.08660

Task fit on one bounded task Two horizontal bars measured against the same bounded task, such as classification. The larger general model used with a prompt alone reaches most of the way. The small model trained and tested on your examples reaches slightly further. A note marks that the comparison holds for narrow, repeated, measurable work. ONE BOUNDED TASK e.g. classification LARGE GENERAL MODEL, PROMPT ALONE SMALL MODEL, TRAINED AND TESTED ON YOUR EXAMPLES holds when the job is narrow, repeated, and measurable
Task fit, not model size. On a bounded task, training and testing on your own examples can matter more than general capability.

The hard part is no longer simply access to computing power. It is deciding what a good answer looks like, collecting representative examples, and building a test that the model must pass before it reaches a customer, an employee, or a consequential workflow.

This matters in Australia for another reason.

From 10 December 2026, certain organisations covered by the Privacy Act must include information in their privacy policies when a computer program uses personal information to make, or substantially assist with, a decision that could significantly affect an individual's rights or interests. Source: oaic.gov.au

The requirement is a disclosure obligation. It does not require a business to own its model. Nor does an open-weight model automatically create an explanation.

But a business that controls a decision system can maintain a clearer record of what it does: the data used, the model and configuration, the evaluation results, the deployment date, and the people accountable for exceptions. That is useful for compliance. More importantly, it is useful when a customer, employee, regulator, or board member asks a simple question: how did this decision get made?

The record a controlled decision system keeps A file card labelled decision record lists five fields: the data used, the model and configuration, the evaluation results, the deployment date, and the people accountable for exceptions. An arrow leads from the card to the question: how did this decision get made? DECISION_RECORD the data used the model and configuration the evaluation results the deployment date the people accountable for exceptions "How did this decision get made?"
Control the decision system and the answer is a record you hold, not a request to a supplier.

Here, ownership does not mean training a model from scratch or insisting that no data ever leaves your premises. It means having enough control to govern the decision rather than merely consume an output. Thinking Machines' Tinker platform, for example, lets users download weights from a trained model for use outside the platform. Source: tinker-docs.thinkingmachines.ai

Start with the highest-volume or highest-risk automated judgment in your business. Then ask four questions.

  1. Is the task narrow and stable enough that we can describe a good answer clearly?
  2. Can our subject-matter experts produce representative examples and an independent test set?
  3. Is the decision frequent or important enough to justify maintaining a dedicated system?
  4. Can we meet the privacy, security, deployment, and human-oversight requirements for running it responsibly?

Four yeses do not mean "build immediately." They mean "run a controlled pilot."

Four yeses mean a controlled pilot Four checks, one per question: narrow and stable, examples and a test set, frequent or important, run responsibly. They feed a gate labelled controlled pilot. Inside the pilot, the specialised system and the frontier model are compared on accuracy, cost, latency, failure modes, and operating load. The outcome line reads: choose the better business outcome. narrow and stable examples and a test set frequent or important run responsibly CONTROLLED PILOT not "build immediately" SPECIALISED SYSTEM FRONTIER MODEL cases neither has seen cases neither has seen accuracy · cost · latency · failure modes · work to operate choose the better business outcome
Four yeses do not mean build. They mean test both options on cases neither has seen, and let the outcome decide.

Test the specialised system against the frontier model on cases neither has seen. Compare accuracy, cost, latency, failure modes, and the work required to operate each option. Then choose the system that produces the better business outcome.

Most of your AI should remain rented. But some of the decisions that make your business distinctive should not live only inside someone else's service.

Rent breadth. Own the decision.

Wes Fischer, founder, NTWRK

Find the decision you should control.

NTWRK builds decision systems you govern: the examples, the evaluation, the record of why, in your environment. The four questions are where every conversation starts.