Human Review Is Not a Scaling Strategy
An AI system can generate work faster than an organization can responsibly approve it. Leaders who promise human oversight need to fund and test the capacity behind that promise.
Commentary
Apple silicon and unified memory made meaningful local model deployment practical inside a personal computer. That change enabled OOMU, and deserves a larger place in enterprise AI strategy.
The assumption that useful AI must arrive through a remote service has a consequence: every new workflow starts with a question about which provider will process the work.
Apple silicon gives us another place to start. For suitable tasks, the computer on the desk can run the model. That changes the product we can build, the dependencies a customer must accept and the economics of repeated use.
This is what enabled OOMU to be built. Apple’s architecture made local inference a practical foundation for the product, rather than an interesting experiment that required customers to become infrastructure specialists.
My argument is that Apple deserves credit for making capable local AI an application opportunity. Enterprise leaders should take that opportunity seriously when deciding which work needs a hosted service and which work can remain on an endpoint.
A language model needs room for its weights, working data and the state associated with processing a conversation. Having enough memory somewhere in a computer does not automatically mean that the processor doing the inference can use it efficiently.
In a conventional discrete-GPU arrangement, system memory and the GPU’s onboard memory are separate physical resources. The model’s execution has to account for that division, whether through transfers, offloading or a different deployment arrangement. Software-managed unified addressing can simplify programming without making those separate resources physically identical. NVIDIA’s CUDA memory guide.
Apple silicon places the CPU and GPU on a shared memory architecture. The MLX unified-memory documentation explains that both processors can access the same memory pool and operate on MLX arrays without moving them between separate CPU and GPU memory locations.
That is a consequential simplification. It allows software to use a substantial shared pool for model execution within a machine that is also the user’s everyday computer. It reduces a class of data-movement work and avoids making a separate graphics card’s memory capacity the sole local constraint.
The pool is not unlimited or entirely available to the model. The operating system and other applications share it. But the architecture changes what developers can attempt without asking the customer to assemble a specialized machine.
Model deployment needs both enough room to operate and acceptable performance. A model that loads but responds too slowly for the task has not solved the product problem.
Apple has combined unified memory with substantial bandwidth in its higher-end chips. Its October 2024 M4 Max announcement specified configurations supporting up to 128 GB of unified memory and up to 546 GB/s of bandwidth. Those are manufacturer specifications for particular configurations, not a promise of inference speed or current availability. Apple’s M4 Pro and M4 Max announcement.
The strategic significance is the combination: accessible capacity, data movement and usable compute in an integrated computer. The result depends on the model, its numerical representation, the inference software, context length and the rest of the workload. A memory-bandwidth figure cannot be converted directly into an application’s response time.
Reducing the precision used to store model weights can make local deployment more feasible, but it still needs a quality test for the intended task. Likewise, a long conversation can require additional working memory. Buyers should evaluate the actual model and workflow while ordinary applications remain open.
The practical question is whether the machine can deliver an acceptable result under everyday conditions. A successful model-loading screenshot is a starting point.

Original Eldris schematic based on the documented shared-memory architecture. Memory capacity, bandwidth and software still determine what a particular machine can run well.
For suitable work on an existing Mac, local inference can avoid a separate GPU server, its administration and a hosted inference charge for each request. It can also avoid making remote inference availability part of the task’s critical path.
That matters for recurring work. Once the hardware and software are in place, additional local inference consumes endpoint resources rather than automatically creating another provider inference bill. A team can experiment without treating each attempt as another hosted transaction.
The costs remain real: hardware, electricity, software, model storage, support and employee time. A new high-memory Mac needs a business case. An existing Mac still has limited capacity, and occupying it has an opportunity cost. Hosted services can be more economical for occasional work or offer capabilities a local model cannot match.
The right comparison includes the whole deployment. An endpoint application, a dedicated local workstation, a managed GPU service and a hosted model API bring different support requirements and capacity assumptions. Comparing a Mac’s purchase price with an API’s unit price settles very little on its own.
Local execution also provides a useful data boundary: information used for local inference need not be sent to an external inference provider. That boundary must be verified across the entire workflow. Connectors, optional cloud actions, telemetry and other services can still involve network activity. Local model execution is one part of a privacy design.
For OOMU, the significance is that we could build around the customer’s machine as a place where AI work actually happens. That is an account of our product’s foundation, not a comparative benchmark or a quantified claim about customer savings.
The hardware does not make the application useful by itself. Software still has to make model selection, resource use and the handling of work understandable. Customers should not have to become experts in GPU memory management before they can benefit.
Other platforms deserve serious consideration. NVIDIA’s DGX Spark, for example, also offers coherent unified system memory. Apple does not own the concept of shared memory, and a purpose-built system may fit a workload better. NVIDIA’s DGX Spark description.
Apple’s contribution to OOMU was the particular combination of architecture and a personal computing platform we could build on. The business claim is strongest at that level: local AI could become an application people use on a Mac, with fewer infrastructure decisions standing between installation and useful work.
For enterprise leaders, the next step is to identify a bounded workload, test it on representative hardware and compare quality, responsiveness, resource use and total operating costs with a credible alternative. Keep the work local where that choice delivers value. Use a hosted service where its capabilities justify the dependency.
Apple silicon helped make that choice real. An AI strategy that assumes every useful model must live elsewhere is leaving an important part of the computer, and the business case, out of the decision.
More insights
An AI system can generate work faster than an organization can responsibly approve it. Leaders who promise human oversight need to fund and test the capacity behind that promise.
A provider's attractive inference price can conceal an expensive dependency. Enterprise buyers should make the ability to change providers part of the original investment decision.
An assistant that presents itself as working for the customer should make commercial influence visible and let the customer define what a good purchase means.
OOMU is free, including at work. If you connect a cloud provider, that provider bills you for what you use.