Skip to content
Back to blog

Private AI Agents: What They Are and How to Choose One

Four boundaries, each scoped on its own termsFour independent horizontal tracks, one each for knowledge, tools, network access, and retention. Each track shows how much of that boundary's own possible range is currently open, scoped on its own terms rather than against a shared scale. One track is drawn open far wider than the other three, showing that an agent's actual reach is set by whichever boundary is loosest, no matter how tightly the other three are drawn.KNOWLEDGETOOLSNETWORKRETENTION
The four boundaries run on different mechanisms and don't substitute for each other: each adds its own exposure, so the loosest one sets what an agent can actually reach, however tightly the other three are drawn. This is a general framework for evaluating any agent platform, not a diagram of Commt's own architecture.

Many people use "private AI agent" to mean an agent that runs in their own environment, on their own servers or in their own cloud account, with data kept inside that infrastructure. That is one useful meaning and the types of private AI agent below cover it. The definition used in this article is wider.

A private AI agent is one whose knowledge, tools, network access and data retention are each deliberately scoped by the organization running it. That's different from leaving each to whatever a platform or model provider defaults to. An agent earns the label "private" only when all four of these are controlled on purpose, not because a vendor states somewhere that its model doesn't train on customer data.

In short: self-hosting an agent, or choosing a model that doesn't train on customer data, can help. Neither alone makes an agent private; both matter, but what actually makes it private is controlling all four boundaries on purpose.

Below: what each boundary means, how to control what an AI agent can access, the types of private AI agent, example tools for each type, how to choose between them, why neither self-hosting nor a no-training clause is enough on its own, and what to ask a vendor before trusting it with anything sensitive.

What actually makes an agent private and what doesn't

Being private is a property of what an agent can reach and what happens to what passes through it. It isn't a property of where it sits, or what a vendor's homepage claims.

What doesn't make an agent private

A few things do not, by themselves, make an agent private:

  • A privacy policy promising good behaviour. Policies describe intent, not what the software can technically do.
  • A "we don't train on your data" line from the model provider. One link in a longer chain, more on that below.
  • Running on your own servers. Self-hosted software can log everything to a world-readable bucket just as easily as hosted software can.
  • The word "private" sitting in the product name.

What actually determines it

What determines it is four boundaries, what Commt calls an agent's reach, each controllable on its own. That's what the agent knows, what it can do, what it can reach over the network and what gets kept afterward. A platform can get any one of these wrong while nailing the other three, so all four need checking, not just the one a vendor happens to lead with.

What are the four boundaries that define a private AI agent?

Four things, each worth checking independently, because getting three right and one wrong still leaves a hole.

Knowledge. Which documents, databases, or knowledge bases can the agent actually read? A private agent's knowledge is limited to what's explicitly bound to it, agent by agent, rather than a shared pool everyone on the account can query by default. Retrieval-augmented generation (RAG) is one way to give an agent knowledge, but it's optional. An agent with no knowledge base attached can still be private, depending on the other three boundaries.

Tools. What can the agent do, beyond generating text? Every connector, API, or function call available to an agent is a door out of the conversation. It's often a door into another system too: a CRM, a ticket queue, a shared drive. Many of these tools reach an agent through MCP, the Model Context Protocol. MCP is an open standard that lets an agent discover and call tools and services, with no custom integration needed for each one.

A private agent's tool list is a scoped allowlist someone deliberately chose, not "everything the platform happens to support." Ask what's on that list and who set it.

Network. Where can requests originate from the agent and where can they go? This is the boundary people get wrong most often, usually by assuming that turning off browsing turns off the network entirely. An agent whose toolset lacks a general-purpose web-browsing or web-search tool cannot fetch arbitrary URLs on its own initiative.

That's different from having zero outbound calls: a connector to Slack or Jira does make an outbound call, on purpose, to the one system it was wired to reach. The useful version of this boundary is "the agent can only talk to what it was told to talk to," not "the agent can't talk to anything."

Retention. What gets kept after the conversation ends and for how long? This covers message content, but also the less visible stuff: request logs, execution traces, tool-call records and anything an observability layer captures on the side. Execution traces are the step-by-step record of what an agent did, including its tool calls.

A private agent's retention is something the customer sets, rather than something the platform defaults to "forever, for troubleshooting." Deletion should mean the record is gone or irreversibly anonymised, rather than soft-deleted and still sitting in a backup someone can restore.

Private AI agents vs public AI assistants

A private AI agent and a general-purpose AI assistant differ in how deliberately their access is scoped, not in which one happens to run in the cloud.

DimensionPrivate AI agentGeneral AI assistant
KnowledgeCustomer-controlledUsually broad and general
ToolsExplicitly configuredUsually platform-defined
Network accessControlled by configurationDepends on product
RetentionCustomer-defined where supportedUsually provider-defined
DeploymentBusiness workflowsGeneral-purpose interaction
Access controlAgent-specificUsually account or product level

A private agent can run in the cloud, browse the web and call outside services. What makes it private is that each of those is a deliberate, checkable setting rather than an assumption inherited from a platform's defaults. That's true whether or not it's self-hosted or cut off from the network.

What types of private AI agents are there?

There are three common types of private AI agent: self-hosted, private cloud and managed platform. They differ in where the agent runs and who maintains it, and none of the three decides by itself what the agent can reach.

"Private cloud" here means the agent runs in a cloud account or virtual private cloud (VPC, an isolated network inside a public cloud provider) that your organization owns and controls.

Self-hostedPrivate cloudManaged platform
Where it runsYour own servers or data centreYour own cloud account or VPCThe vendor's hosted infrastructure
Who maintains itYour team: model updates, patching, scalingYour team, with the cloud provider running the underlying hardwareThe vendor runs the platform; you configure each agent
Boundaries you controlAll four in principle, limited by what the software offers and what your team buildsAll four in principle: the cloud account gives you network and storage control, the agent software decides the restAll four, within what the platform exposes
FitsOrganisations with a custody or residency requirement and a team to run the softwareTeams already operating in their own cloud that need data to stay in their own accountTeams without capacity to run an agent platform
Main trade-offYour team carries the operational load: updates, patching, scaling and loggingYou still own deployment, patching and monitoringThe platform runs on the vendor's infrastructure, not yours

Where an agent runs is one input. It decides who has custody of the infrastructure. It does not decide what the agent can reach. Two self-hosted agents can differ completely: one with a scoped tool allowlist and no stored messages, another with open browsing and full logs. Whichever type you pick, check the four boundaries on the specific agent.

A managed platform such as Commt is not self-hosted. Commt's application, database and document storage are hosted in the EU on private networking, and at launch the customer controls each agent's knowledge, tools and network access, as well as data retention. If your requirement is that the whole workload runs inside infrastructure you control, a managed platform does not meet it and one of the first two types does.

What are some private AI agent tools?

Tools that fit the three types include OpenClaw, Hermes Agent, n8n and Onyx (self-hosted), Amazon Bedrock AgentCore, Microsoft Foundry Agent Service and Gemini Enterprise Agent Platform (private cloud), and Dust and Microsoft Copilot Studio (managed platform).

These are examples, not a ranking or an endorsement, based on each vendor's own documentation as of October 2026. The table shows where each one runs. The four boundaries still have to be checked on each agent you deploy.

ToolTypeWhat it isWhere it runsLicense or terms
OpenClawSelf-hostedOpen-source personal AI assistant that works through chat appsYour own hardwareMIT
Hermes AgentSelf-hostedOpen-source agent from Nous ResearchLocally or in Docker, with hosted optionsMIT
n8nSelf-hostedPlatform for building AI agents and workflowsSelf-hosted or on n8n's cloudSustainable Use License (n8n calls it "fair-code"; not OSI-approved open source)
OnyxSelf-hostedKnowledge layer, formerly DanswerDocker or Kubernetes, including air-gapped, or Onyx CloudMIT for the Community Edition, with a separate Enterprise Edition
Amazon Bedrock AgentCorePrivate cloudAWS agent runtimeRuntime attaches to your VPC; in that setup, no internet access by defaultCommercial service
Microsoft Foundry Agent ServicePrivate cloudMicrosoft's service for building agentsWith private networking, agent data stays in your Azure tenantCommercial service
Gemini Enterprise Agent PlatformPrivate cloudGoogle's agent platform, previously Vertex AIPrivate VPC connection; internet access must be configured explicitlyCommercial service
DustManaged platformHosted platform for building agentsHosted by Dust, with US or EU data residencyCommercial service
Microsoft Copilot StudioManaged platformMicrosoft's studio for building agentsAdmin data policies can block knowledge sources, connectors and channelsCommercial service
CommtManaged platformAgent platform; at launch, customers set each agent's knowledge, tools and network accessHosted in the EU on private networking; not self-hostedCommercial service

The three private cloud services are run by the cloud provider and attach to your network, so your team maintains the deployment around them, not the service itself. Choose by which type fits, then check the four boundaries on the agent itself.

Is self-hosting the same as private?

Not the same thing. Self-hosting is one way to get some of these privacy properties and for a real subset of buyers, it's the right way. But delivering the four boundaries above still takes deliberate configuration on top of it, whichever way you host.

Self-hosting wins when the requirement is about who has physical or administrative custody of the infrastructure. That covers a government agency, a defence contractor, or a bank operating under a rule that the workload has to run inside a boundary the vendor doesn't control. If that's the actual requirement, no amount of vendor promises about scoping and retention substitutes for it.

The cost side of self-hosting, for most buyers

That's a different calculation from the one facing buyers who lack a custody requirement, which is most of them. For that group, what gets underweighted is the cost on the other side. Self-hosting an agent platform means your team now owns model updates, security patching and scaling under load. It also means building the same knowledge, tool, network and retention controls a hosted vendor would otherwise maintain. Getting those four right is the hard part, regardless of who runs the servers.

A self-hosted deployment only stays private if someone actually staffs it, keeping its logging tight and its connector scope narrow. Skip that and it's just as exposed as a badly run hosted setup, misconfigured somewhere your own team can't see. Most teams lack a standing platform security function. For them, a hosted vendor that publishes and enforces the four boundaries above will beat a self-hosted setup with no one dedicated to hardening it.

Does "the model doesn't train on my data" mean the agent is private?

Not by itself. A no-training commitment from a model provider covers exactly one hop: the moment your prompt reaches that specific model. It says nothing about everything else that happens to your data on the way there and back.

What sits between your message and the model

Between a user's message and a model's response there's usually a platform layer.

  • Request logs and whatever the agent's tools returned.
  • An orchestration or agent framework, the layer that sequences an agent's steps and tool calls, which may retain transcripts for debugging.
  • Analytics or tracing systems capturing payloads for observability, meaning visibility into what the system did, used for debugging and monitoring.
  • Any subprocessors sitting underneath, the other companies the vendor relies on to run parts of its service.

A no-training clause is silent on all of it. A platform can honestly say the model provider doesn't train on your data, while still storing full conversation transcripts indefinitely in its own database. It might also ship tool-call payloads to a third-party logging service, or give a wide internal team read access to customer traces.

Privacy is a property of every hop, not the last one.

Treat a no-training commitment as answering one question, not the whole audit. Private AI agent vs. ChatGPT applies that test to ChatGPT Business.

9 questions to ask before choosing a private AI agent platform

Ask questions specific enough that a vendor without the right architecture can't dodge them.

  • Which knowledge bases can this specific agent read and can I see that list, not just the account-wide one?
  • What tools or connectors can the agent invoke and who chose that scope, me or your default?
  • Is there a browsing or web-search tool available to agents on your platform and if not, is that a setting or an architectural fact?
  • Your connectors make outbound calls; where exactly do those calls go and is that list fixed or something the agent decides at runtime?
  • What message content do you store by default and for how long, before I change anything?
  • If I delete a conversation, is the record gone, retained and anonymised, or retained and just hidden from my dashboard?
  • Where do request logs and execution traces live and who at your company can read them?
  • Does your "we don't train on customer data" statement cover your own platform logging and your subprocessors, or only the underlying model?
  • If I self-host instead, what operational burden moves to my team and have you sized it honestly?

A vendor with a genuinely private architecture can answer all nine specifically. A vendor selling "private" as a feature usually stalls around the third one.

Are private AI agents offline?

Not necessarily. A private agent can run in the cloud and still have tightly controlled access to knowledge, tools, networks and retained data. Privacy is about controlling what an agent can reach and what happens to its data, not whether it's connected to anything. A fully disconnected agent is one specific configuration, not a requirement for being private. What matters is that its network access, however much or little, was set on purpose, rather than left to a platform's default.

Are private AI agents only for enterprises?

Not just enterprises. Any business that needs control over what an AI system can reach benefits, including customer support, sales, operations and internal knowledge workflows. How much control is needed depends on the data and systems the agent touches, rather than the size of the company running it.

Key takeaways

  • Private is defined by controlled access, not by where an agent runs.
  • Private AI agents come in three types: self-hosted, private cloud and managed platform. Where they run changes who maintains them, not what they can reach.
  • Tools exist for each type, but a tool's type tells you where it runs, not what it can reach.
  • Knowledge should be scoped explicitly, agent by agent, rather than shared across an account by default.
  • Tools and connectors need an explicit allowlist someone chose on purpose, not whatever a platform happens to ship.
  • Network access is a deliberate, configurable setting, not a single blanket switch.
  • Retention should be a policy the customer sets, with deletion meaning the record is gone or irreversibly anonymised.
  • Self-hosting can help with a genuine custody requirement, but it doesn't automatically make an agent private.
  • A no-training policy from a model provider covers one hop in the chain, not the whole path your data takes.

What does this look like in practice?

Commt, an AI agent platform, scopes each of the four boundaries by design:

  • Knowledge. An agent reads only the knowledge bases explicitly bound to it, with RAG optional rather than assumed.
  • Tools. Which tools and connectors an agent may reach is set by the customer, from a list of twenty connectors available at launch.
  • Network. Web browsing and web search are available to agents and customers can disable them when creating an agent.
  • Retention. No message content is stored by default, with retention and deletion set by the customer, meaning the record is deleted or anonymised.

Commt is a managed platform and is not self-hosted. It's one example of applying the checklist above, not the only correct shape a private agent can take.