Self-Hosted vs Cloud AI Agents, Which One Passes a Production Security Review?
Compare self-hosted and cloud AI agents that read production, with five deployment shapes, eight security review questions, and where LLM prompts go.
Quick Answer
A self hosted AI agent and a cloud AI agent differ by where the agent runs, where its production credentials live, and which LLM receives log lines, rows and code. Security reviews move fastest when every party that receives data is named. Vendors sell five deployment shapes, from multi-tenant SaaS to full self-hosting with an open-weight model. Self-hosting keeps data home only if the model also runs at home.
Self-hosting means running software on servers or cloud accounts you control, instead of on the vendor's infrastructure. For most tools that is a hosting preference. For an AI agent that reads your production logs, your Postgres replica and your repo, it decides who can see your customers' data.
The question usually lands the same way. An engineering leader wants faster root cause, and security wants to know where the logs go. The vendor answers "we're SOC 2" and the review stalls, because nobody drew the data path.
In this post we'll cover the five deployment shapes vendors actually sell and the security review questions that separate them. We'll also cover the LLM layer, what self-hosting costs a small team, and a five-step way to decide in one meeting.
Self-Hosting or Cloud for an AI Agent That Reads Production, the Short Answer
Choose self-hosting when your logs or database rows contain regulated or contract-restricted data and you can name someone to run the agent. Choose a cloud agent when your data can be redacted, the vendor names every subprocessor, and you have no one to patch another service.
The mistake we'd warn against first is treating this as a two-way choice. A review is really about three separate boundaries, and each one can sit inside or outside your network.
| Boundary | What it controls | Question to ask |
|---|---|---|
| Agent compute | Where the investigation code runs | Does the agent run in your account or the vendor's? |
| Credential custody | Where the read-only keys and database passwords are stored | If the vendor is breached, do your production credentials leak? |
| Model inference | Which large language model receives prompts full of your data | Which company receives log lines, and how long do they keep them? |
A software as a service (SaaS) agent usually puts all three outside. Full self-hosting with an open-weight model puts all three inside. Everything else is a mix, and the on prem AI vs cloud debate only gets useful once you say which boundary you mean.
Pro tip: Before any vendor call, we write the three boundaries on one slide with "inside" or "outside" next to each. It turns a vague privacy debate into three yes-or-no answers.
Five Deployment Shapes AI SRE Vendors Actually Sell
Vendors describe at least five shapes, and self-hosting is only the last of them. We read the security pages of four AI SRE vendors on 29 September 2026, and their own wording maps onto these shapes.
1. Multi-Tenant SaaS That Holds Your Credentials
The agent runs in the vendor's cloud next to other customers. You paste read-only API keys and a database connection string into their UI, and their servers connect to your systems. That usually means allowlisting vendor IP ranges or exposing an endpoint, and your credentials now live in their secrets store.
2. SaaS With a Connector Inside Your Network
A small component runs inside your environment and forwards data out on request. Resolve AI's security page calls this a "satellite" that can proxy data from your tools or read only Kubernetes metadata. Credentials can stay local, but the evidence an investigation needs still leaves.
3. Single-Tenant Dedicated Cloud
The vendor runs an isolated instance just for you. Resolve AI lists a dedicated single-tenant option with bring-your-own-keys encryption. This removes cross-customer risk, yet the data still sits in an account you do not own.
4. Bring Your Own Cloud
The vendor's software runs inside your cloud account, often inside your virtual private cloud (VPC). Traversal's security page offers bring your own cloud and private on-premises deployments. We always ask what the vendor's control plane can still see, because "runs in your account" and "phones home with nothing" are different claims.
5. Full Self-Hosting With Your Choice of LLM
The agent runs as a container in your infrastructure, stores credentials in your database, and calls whichever model you configure. This is the only shape where agent compute and credential custody are both inside by default. Model inference is still a separate choice, because the agent can call a hosted API or a model on your own hardware.
Here is what each vendor publishes about its own options, as checked on 29 September 2026:
| Vendor | Deployment options on its own pages | Data and model handling it states |
|---|---|---|
| Resolve AI | Shared multi-tenant with a satellite and dedicated single-tenant cloud | No write access, least privilege, no storage of raw data, SOC 2 Type II |
| Traversal | Private on-premises; bring your own cloud | Read-only access; bring your own model, including self-hosted; SOC 2 Type II |
| incident.io Investigations | Hosted service with a per-customer instance of its Nexus model | Redacts sensitive data before model providers, zero data retention agreements with them, and only change is a pull request you merge |
| Datadog Bits Investigation | Runs inside Datadog; not supported on the app.ddog-gov.com and us2.ddog-gov.com sites | The overview page we read does not describe model data handling |
None of these is wrong. Each one answers a different threat model, and we'd pick the shape before we pick the vendor.
The Security Review Questions That Separate Self-Hosting From Cloud
We reduce an AI agent security review to eight questions, and six of them get different answers under self-hosting and cloud. The other two, write access and prompt injection, apply to every shape equally. This table is the map we use:
| Review question | What a cloud agent must show | What self-hosting must show |
|---|---|---|
| Where do the read credentials live? | Vendor secrets store, encryption, rotation, breach scope | Your secrets store and who can read it |
| Which new network path opens? | Inbound from vendor IPs, or an outbound connector | Usually only outbound to the model endpoint, if hosted |
| What leaves per investigation? | Log lines, query rows and code sent to the vendor and its LLM provider | Only what you send to your model provider, if any |
| Who are the subprocessors? | The vendor's list, including every model provider | Your model provider, under your own contract |
| Where are findings stored? | Vendor database, with its retention period | Your database, with your retention period |
| What evidence of controls exists? | A SOC 2 Type II report and pen test summary | Your own controls, image scanning and audit log |
| What can the agent write? | Same for both shapes | Same for both shapes |
| Can a log line steer the agent? | Same for both shapes | Same for both shapes |
• Credentials and Network Paths
Credential custody is the row reviewers underrate. A read-only role still reads every customer record it can reach, as we argued in read-only is a permission, not a capacity limit. With a cloud agent, that role's password sits in someone else's breach scope.
• Data Egress and Subprocessors
Under GDPR Article 28(2), a processor may not engage another processor without the controller's written authorization. An AI vendor that sends your log excerpts to a model provider is doing exactly that. We ask for the subprocessor list on the first call, not the last.
• Agency, Prompt Injection and Audit
OWASP's LLM06:2025 Excessive Agency names excessive functionality, permissions and autonomy as three root causes. A log line is untrusted input, so prompt injection can arrive inside the very data the agent reads.
We want the agent's only output to be a draft a human reviews, because a pull request is write access with a waiting period. We also want every step logged, as in AI-assisted infra changes need a paper trail.
This is where the deployment model stops mattering. A self-hosted AI agent with write access is riskier than a SaaS agent held to read-only roles and draft-only output.
Self-Hosting the Agent Does Not Self-Host the LLM
Self-hosting the agent keeps credentials and findings home, but prompts still go wherever the model runs. The self hosted LLM vs OpenAI question is a separate decision, and it carries its own retention terms. We checked each provider's own documentation on 29 September 2026:
| Model path | Who receives the prompt | Documented retention |
|---|---|---|
| OpenAI API direct | OpenAI | Up to 30 days to provide the service and identify abuse; zero data retention for eligible endpoints on request |
| Anthropic API direct | Anthropic | Deleted within 30 days, except under a zero data retention agreement, usage policy enforcement or law |
| Amazon Bedrock in your AWS account | AWS, and model providers have no access to Bedrock logs or prompts | Per-Region mode can be none, default or aws_review |
| Open-weight model you serve with vLLM | Nobody outside your network | Whatever you configure |
• Hosted API Direct
This is the simplest path and the one we'd expect a pilot to start on. Your contract with OpenAI or Anthropic becomes the relevant data processing agreement, so the review question moves from "the agent vendor" to "the model vendor."
• Cloud Platform Inside Your Account
Bedrock keeps inference inside your AWS relationship, and its data protection page says model providers cannot see customer prompts. The trap is that Bedrock's docs state that Claude Fable 5 and Fable 5.1 require the aws_review mode. That retains prompts for up to 30 days inside AWS, so an account-wide zero data retention policy makes those models unavailable.
• Open-Weight Model on Your Own Hardware
Serving an open-weight model keeps every prompt inside your network. vLLM ships an OpenAI-compatible server, so an agent built for the OpenAI API format can often point at it by changing the base URL. What does not transfer is quality. We'd test the local model on your last ten real incidents before trusting it at 3 a.m.
Pro tip: Ask the agent vendor which model version each investigation step uses. If routing is dynamic, your zero data retention terms have to cover every provider in the pool.
What a Self Hosted AI Agent Costs a Team of 15 to 150 Engineers
Self-hosting is cheap in license terms and expensive in attention. The agent becomes one more production service your team patches, backs up and pages on. In our view, that cost often decides the answer before security does for teams in this range.
• The Agent Is Now a Production Service You Run
Someone upgrades the container image, rotates its credentials, keeps its database healthy and watches its audit log. If that someone does not exist on your team today, a self-hosted AI agent will drift out of date. We'd rather see a well-reviewed cloud agent than an unpatched self-hosted one.
• Model Hosting Is the Expensive Part
Running your own model is where the bill moves. Prem AI's cost comparison, from a vendor that sells self-hosting, estimates about $2,100 a month to lease an A100 for Llama 3.1 70B. For most teams in this range, we'd self-host the agent and call a hosted model under a zero data retention agreement instead.
• When a Cloud Agent Is the Better Call
A cloud agent wins when your logs carry no regulated data, the vendor shows a SOC 2 Type II report, names its model subprocessors and redacts before inference. It also wins when nobody owns another service.
The AI data privacy concerns are real, but so is the risk of a tool nobody maintains, and teams building their own ops copilot run into the same staffing wall.
How to Decide in One Security Review Meeting
You can settle the choice in one meeting by drawing the data path for a single real incident and naming every party that touches it. We run this five-step procedure instead of sending a long AI vendor security questionnaire first.
1. Trace One Real Investigation End to End
Pick last month's worst incident. List exactly what an agent would have read, including which log lines, which tables and which files in the repo. Say your checkout logs include customer emails; that single fact already rules out any shape that ships raw logs to a party without a data processing agreement.
2. Name Every Party That Receives Bytes
Write down the agent vendor, the model provider, any cloud platform in between, and where findings are stored. If a vendor cannot name its model provider, the meeting is over.
3. Match Each Party to Your Data Classification
Regulated data that must stay in your network points to self-hosting with a local or in-account model. Internal-only data usually passes with a hosted model under zero data retention terms.
4. Pick the Least Exposed Shape Your Team Can Operate
Choose the shape with the fewest outside parties that still has an owner on your team. Bring your own cloud and a dedicated tenant are real middle options, not compromises.
5. Put Agency Limits in the Database, Not the Prompt
Whatever the shape, enforce read-only access where the data lives, following the principle of least privilege. For a Postgres replica we start with a role like this.
CREATE ROLE ai_agent_ro LOGIN;
GRANT pg_read_all_data TO ai_agent_ro;
ALTER ROLE ai_agent_ro SET default_transaction_read_only = on;
ALTER ROLE ai_agent_ro SET statement_timeout = '5s';
Two caveats. pg_read_all_data reads every table, so we swap it for grants on specific schemas when PII lives nearby. And default_transaction_read_only is only a default a session can change; the real guard is that the role holds no write privileges.
We cover the rest of this pattern in automating root cause analysis without giving AI write access.
Pro tip:
statement_timeoutdefaults to 0, which means no limit. An agent that writes a bad join on a busy replica is a very efficient way to start a second incident.
Pick the Shape Your Security Team Can Draw on One Page
Self-hosting remains the core issue, but self hosted AI vs cloud is really three choices, where the agent runs, where credentials live, and where the model runs. We'd draw one real investigation, name every party that receives data, and pick the least exposed shape your team can actually operate.
A self hosted AI agent paired with a hosted model is often the practical middle, as long as the model contract carries zero data retention.
If you want to see that shape in practice, Operate runs as a Docker image in your infrastructure with read-only access. It uses the AI provider and keys you choose, and drafts fixes for an engineer to approve.
Frequently Asked Questions
Yes. We'd point teams at open-weight models such as Llama, Qwen or Mistral, whose licenses allow commercial use, served with a tool like vLLM on their own GPUs. Self-hosting the model means the GPU must hold the full weights in memory at the precision you choose.
Not as model weights on your own servers, as far as we found. Anthropic's documentation lists the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud's Agent Platform and Microsoft Foundry as the routes to Claude. For self-hosting purposes, the closest option we know is running it inside your own cloud account.
It depends on traffic more than hardware. Hosted APIs bill per token, so light use costs little, while a self-hosted model costs the same at one request or ten thousand. We'd consider self-hosting a model only once usage is steady and high enough to keep the GPU busy.
For operations work we almost never build a model at all. Training from scratch is out of reach for most teams, so the real options are prompting a pretrained model or fine-tuning one. Self-hosting only matters for the fine-tuned copy, which you can then serve locally or in your cloud account.
Privacy depends on retention terms, not brand names. We check how long a provider keeps flagged content. Anthropic's Privacy Center says flagged inputs and outputs are kept up to 2 years and feedback submissions for 5 years. With self-hosting on your own hardware, we set the retention policy ourselves.


