The Fourth Circle on my work table

A local LLM on the workbench connected to a hybrid cloud ecosystem

Date of First publication:

|

Updated on

From rented intelligence to hybrid cognitive ecosystem: why a local LLM does not replace the cloud, but gives us back the possibility of choosing.

When I wrote in 2024Exploring Cloud-Native Ecosystemsand introduced the paradigm ofFourth Circle, I was trying to describe a new dimension of the ecosystem: an intelligence capable of crossing infrastructures, platforms, applications and organizations, changing the way we make decisions and produce value.

Then I imagined above all an AI inserted into the cycle of human activities: oneAI in the loop. Only two years later we discussHuman in the Loop, almost taking it for granted that the cycle now belongs to the machine and that our problem is to preserve a significant space within it for the human being.

In the meantime, another movement occurred, less flashy but equally important. For years we have been bringing data, applications and computing power to the cloud. Now some intelligence is making the reverse journey: returning close to data, processes and people. It can be performed on a compact, silent machine, literally placed on our work table.

The Fourth Circle, born as a dimension of the cloud-native ecosystem, can today also inhabit a small infrastructure that we own and govern directly.

Two points of view that consolidated my vision

The technical demonstrations ofAlex Ziskindand the reflections ofManolo Remiddithey helped me make more concrete a vision that was already, by its nature, holistic.

Ziskind shows the operational side: on Apple Silicon you can run open source models through tools like Ollama, MLX and Open WebUI, building a local service compatible with applications and agents on your hardware. Apple itself now presents a stack forAgentic AI run entirely on the Mac: local model, server compatible with OpenAI APIs, tools and agent cycle.

Remiddi introduces a deeper provocation: the difference betweenrented intelligenceand cognitive ability possessed. In his journey on theconstruction of a second local brain, hardware is not just a technical choice: it becomes a form of sovereignty over data, memory and one's cognitive environment.

But Remiddi himself warns thatlocal AI alone is not enough. Running a model on your computer changes where inference occurs, but it does not automatically address where the training data comes from, the transparency of the model, the quality of the responses, or control over its evolution.

This is where I find my perspective: we don't have to ideologically choose between cloud and on-premise. We have to design ahybrid cognitive ecosystem, deciding which intelligence, which data and which action belong to each perimeter.

Because this change is happening right now

The first reason is the evolution of open source models. China is not competing only through investment, but through open, efficient models that are progressively closer to the performance of Western proprietary systems. TheStanford AI Index 2026notes that the performance gap between the best US and Chinese models has essentially closed, with a distance of a few percentage points in the evaluations considered.

Families like Qwen, DeepSeek, and MiniMax make it possible to run capabilities locally that until recently required remote services.Reuters documented ithow demand for low-cost, open source-based Chinese solutions is growing rapidly. AlsoDeepSeek continues to push efficiency, reduced costs and technological autonomy; at the same time Western companies begin to experiment with models derived from Qwen to reduce dependence on proprietary services.

As software becomes more accessible, however, the scarcity shifts. The real constraint of local inference is not just the processor: it is above all the availability offast memoryin which to load models and context.

Demand from AI data centers is already putting pressure on the entire memory chain. Reuters reported aglobal memory chip supply crisis, with forecasts of significant increases, and subsequently linked the increase in prices of smartphones and PCs precisely to the demand coming from the AI ​​infrastructure. It is therefore not unreasonable to hypothesize that configurations best suited to local inference — lots of RAM or unified memory, high bandwidth, and low power — may become more popular and expensive.

The other side: the cost of intelligence in the cloud

Even intelligence offered as a service has an economy that is anything but stabilized. Each response requires computing capacity, memory, energy and network infrastructure. The more users, contexts and agents increase, the more the number of inference cycles needed grows.

La Stampa reported on OpenAI's efforts to halve the cost of inference: a decisive optimization, because a few cents per million tokens can separate a sustainable service from a loss-making one. The same newspaper highlighted the paradox of a platform with a huge audience, but with a majority of free users who represent a cost even before a revenue.

The pressure isn't just for OpenAI. According to oneReuters analysis, the growth of AI investments is squeezing hyperscalers' cash flow, and investors are demanding more concrete evidence of economic returns. Reuters also estimated the AI ​​investment plans announced by Big Tech for 2026 at around $600 billion.

We cannot know if the price of each individual subscription will increase. However, we can build a reasonable scenario in which platforms introduce premium bands, more stringent limits, agent-specific costs, extended contexts and frontier models. Dependence exclusively on the cloud means remaining exposed to economic and commercial decisions made by others.

The venue as a calmer, not as a substitute

A local host does not have to beat the best existing model in every test. It must reach a much more concrete threshold:do what I or my organization needs reasonably well and consistently.

  • classify and summarize documents;
  • query a private knowledge base;
  • assist in code development and analysis;
  • process data that must not leave the perimeter;
  • run repetitive, high-volume agents;
  • guarantee a minimum capacity even without access to the remote service.

When a configuration reaches this threshold, the release of a more powerful machine does not automatically render it useless. I can continue to use it for years, amortizing the initial investment. The local hardware thus becomes acalming: Does not eliminate the cloud, but limits exposure to future increases in its prices.

The Fourth Circle on my table does not replace the one in the cloud: it gives me back bargaining power over it.

From OPEX only to a CAPEX/OPEX portfolio

The cloud has accustomed us to buying capacity as an operational expense: I pay every month or for each token used. An AI workstation instead introduces an investment component: I purchase a capacity today that remains available over time.

This does not mean that the venue cancels OPEX. That leaves energy, maintenance, storage, management and updates. More correctly, the hybrid architecture allows you torebalance CAPEX and OPEX:

  • basic ability possessed, stable and predictable;
  • advanced ability purchasedwhen complexity and value justify it;
  • task routingbased on cost, confidentiality, latency and required quality.

It's the same economic logic addressed in the articleAI in the SDLC: is your business gaining or losing?: the price of the instrument is not enough. We need to measure the overall value produced or protected.

A formula to calculate how much I earn

Hybrid benefit = cloud costs avoided + value of time saved + value of privacy and continuity − local TCO

Where:

Local TCO = hardware + energy + storage + maintenance + management time

Over a period of a few years, the comparison becomes:

Benefit = cost of cloud-only scenario − (hardware CAPEX + local OPEX + remaining cloud)

The voiceresidual cloudit is essential. The goal is not to isolate ourselves: we will continue to use frontier models for more complex problems, for updated information or for peaks in capacity. A local model will instead be able to absorb frequent, private and repetitive work.

The break-even point changes for each person and organization. If I use AI occasionally, the hardware may not be convenient. If I process large volumes, sensitive data, or continuous tasks, the value of the capacity I own grows rapidly. Privacy, resilience and provider independence are not abstract benefits: they must enter into the calculation with a value consistent with the risk avoided.

At home and in the office: the same circle, different responsibilities

On your home desk, the local LLM can become a form ofPersonal Sovereign AI: memory, documents, writing and experimentation remain within my scope. In the office, however, it becomes a hubEnterprise Edge AI, integrated with identity, authorization, audit, data classification and organizational policies.

Physical proximity does not eliminate governance problems. A local agent with access to files can delete, modify, or disclose information just like a remote agent. We will therefore have to design boundaries, levels of autonomy, tracking of actions, checks and possibilities of arrest.

The Fourth Circle is not sovereign just because it runs on our hardware. It becomes so when we know:

  • where model, data and memory reside;
  • what tools it can use;
  • who authorizes his actions;
  • how results and accountability are verified;
  • when the task needs to move from local to cloud or back to human.

The return journey of intelligence

For years, cloud-native taught us that owning the infrastructure wasn't always necessary. We could consume elastic capabilities, evolve them rapidly, and pay as we go. These advantages remain valid.

But AI introduces a new question: how much of our cognitive capacity do we want to continuously purchase from external platforms and how much do we want to stabilize in our perimeter?

The answer will not be the same for everyone. Mine, today, is not an escape from the cloud. It is the construction of an ecosystem in which cloud, edge, local host, open models and proprietary services cooperate under human direction.

I'm not simply bringing an LLM to my desk. I'm recomposing an entire cognitive ecosystem on my table.

The Fourth Circle got close enough to see him work. Precisely for this reason we must choose with greater awareness where to place it, how much to delegate to it and which part of our autonomy we want to continue to safeguard.

Transparency note
The ideas, vision, narrated experience and editorial responsibility of this article belong to Mauro Giuliano. Faber, digital assistant to the editorial staff of Exploras.cloud, collaborated in the research of the sources, in the structuring and revision of the text under the direction and with the final approval of the author.


References

Comments

Leave a comment

Your email address will not be published. Mandatory fields are marked*

This site uses Akismet to reduce spam.Find out how data derived from comments is processed.