Trust.
Alovia runs in front of your website (Shield) and in front of your AI agents (Watchdog). Both sit in the path of your traffic, so you should be able to see how we run them, what we keep, what leaves our systems, what we measure, and how to tell us when something is wrong. This page states all of that plainly, including the things we have not done yet.
Where a fact below comes from the code that runs, it is read from that code and changes only when the code does. Where it is a certification or an audit, it is marked as done or not done. Nothing here is a claim we cannot point at.
- Hosting
- United States. Application, database and queues on Railway; Shield edge servers on DigitalOcean. See Subprocessors.
- In transit
- TLS on every connection to Alovia. The gateways reach your own host at the address you configured, so that leg is encrypted when the address is https.
- Retention
- Watchdog message excerpts and review rationales cleared after 30 days. Shield traffic history up to 90 days. Verdict records kept for your audit trail until you delete the account.
- Training
- Never used to train generative AI. Improving detection with your redacted excerpts is off unless you turn it on, and turning it off erases what you contributed.
- Certifications
- SOC 2 Type II: not yet. ISO 27001: not yet. Penetration test: internal only, no third-party letter yet.
- Report a vulnerability
- security@aloviaai.com, under the disclosure policy below. Safe harbor applies.
- Status page
- Not yet. Incidents are announced by email to affected accounts.
How we protect your data
Where it runs. Everything runs in the United States. The application, the Postgres database, Redis and the background queues run on Railway. Shield's edge servers, the ones that sit in front of protected websites, run on DigitalOcean. Both providers hold their own attestations; our shared-responsibility slice is described in the Data Processing Addendum.
Keys and secrets. Your fleet key is stored as a hash for lookup and as an encrypted copy under a server-held key, so it can be shown to you again from the dashboard; the deploy key is stored the same way. Our own secrets are provisioned from the secret store as environment variables and never land in source control.
Content is redacted before it is stored. The 160-character excerpts the dashboard shows you pass through a redactor that removes authorization headers, bearer tokens, vendor key prefixes, JWTs and values assigned to sensitive key names. It covers the common shapes and is deliberately not exhaustive; a private key block or a secret with unusual punctuation can survive it.
Tenant isolation is enforced in application code. Every query filters by your account, and every route checks ownership before it reads. The database itself does not distinguish one tenant's rows from another's: row-level security is a deny-all lockdown, not per-tenant policies. We say this because it is the honest description of where the boundary lives, and per-tenant database policies remain on our list.
Who can touch the code. Multi-factor authentication is enforced across the GitHub organization. The main branch requires signed commits and a review; no direct pushes, no forks. Third-party OAuth apps are audited quarterly.
Retention is a number, not a promise. Excerpts and review rationales are cleared by a daily sweep after 30 days. Findings you dismissed or fixed are kept as decisions (detector, severity, your note) without the message. Account deletion removes everything within 30 days, except data under a legal hold.
What leaves our infrastructure
Two paths carry customer-derived content off our systems. Neither is on by default. Nothing else does.
- —The review model. When a detector is uncertain about a message, Watchdog can send that message, the agent's mission and up to 4,000 characters of recent conversation to Anthropic for a second opinion. It fires only on that uncertain slice, never on every message, and only when the review is enabled. Anthropic is therefore a subprocessor for accounts with the review on; with it off, uncertain messages are recorded as unreviewed rather than lost.
- —Your own telemetry collector. If you point Watchdog at an OpenTelemetry endpoint, findings are exported there. The export carries the finding's type, severity, detector, technique ids and a pattern-level detail line. It does not carry the message, the excerpt or the review rationale. The destination is yours.
The gateways call one place. The A2A and MCP gateways forward traffic to exactly one destination: the host you configured, which must be a domain you proved you own. Before every request the address is checked against your verified domains and against private address space, so our egress can never be pointed at anything else by a database row or a caller.
Standards coverage
What Watchdog defends against, in the two vocabularies a security review uses. Both lists are read from the detectors and controls that run. Covered means a detector or control exists for the item; it does not mean every attack in that class is caught. The measured numbers are in the model card below.
OWASP Top 10 for Agentic Applications 2026
| Risk | Detects | Controls |
|---|---|---|
| ASI01 Agent Goal Hijack | encoding, prompt injection, novel intent, unexpected language, mission deviation, obfuscation | Mission scope enforced on every call · Envelope, role-marker and many-shot signals on tool results · Injected span cut out, remainder relayed only if clean · System-prompt fingerprints as a credential-class leakopt-in |
| ASI02 Tool Misuse & Exploitation | exfiltration chain, destructive command, undeclared host, data leak, mission deviation, internal-address access | Tool allow-list, deny unlisted tools · Counterparties bound to the user's request · Response egress gate on both gateways · Hostnames in arguments resolved before the dial |
| ASI03 Agent Identity & Privilege Abuse | out-of-scope reach | OIDC caller identity, fail-closedopt-in · did:web + VC-JWT delegation, anchored on host ownershipopt-in · Delegation can only narrow the mission · AP2 mandate caps: counterparties, uses, amountopt-in · Deploy key and runtime key split |
| ASI04 Agentic Supply Chain Compromise | — | Tool-listing poisoning gate and per-tool digest pin · Signed agent-card verificationopt-in · Upstream must be a host the tenant proved it owns |
| ASI05 Unexpected Code Execution | destructive command | Destructive-hint tools held for a human |
| ASI06 Memory & Context Poisoning | memory poisoning | Memory writes judged, retrieval results checked, provenance ledger |
| ASI07 Insecure Inter-Agent Communication | — | A2A gateway: pinned transport, spoofed headers stripped, no redirects · Signed agent-card verificationopt-in |
| ASI08 Cascading Agent Failures | cascade anomaly | Cascade graph with downstream freeze propagation · Runaway-loop and fleet-wide loop caps · Judge spend bounded per tenant, per tier and per caller |
| ASI09 Human-Agent Trust Exploitation | — | Human approval hold with single-use grantsopt-in · Held-queue flood cap · Session escalation hold after repeated uncertain signals |
| ASI10 Rogue Agents | ungoverned agent | Default-deny for an agent with no mission after a grace window · Agent-scoped auto-quarantine · Canary tool, address and secret per fleetopt-in |
Titles checked against the genai.owasp.org listing, 2026-09-25.
MITRE ATLAS
12 techniques covered, ids checked against ATLAS v5.6.0: AML.T0025 Exfiltration via Cyber Means; AML.T0051 LLM Prompt Injection; AML.T0051.000 LLM Prompt Injection: Direct; AML.T0051.001 LLM Prompt Injection: Indirect; AML.T0053 AI Agent Tool Invocation; AML.T0057 LLM Data Leakage; AML.T0068 LLM Prompt Obfuscation; AML.T0070 RAG Poisoning; AML.T0080 AI Agent Context Poisoning; AML.T0080.000 AI Agent Context Poisoning: Memory; AML.T0086 Exfiltration via AI Agent Tool Invocation; AML.T0101 Data Destruction via AI Agent Tool Invocation. The Navigator layer imports over the ATLAS matrix and names the detector behind each technique.
Model card: the review model and the detectors
Watchdog decides in two layers. Sixteen deterministic detectors (pattern and structure matching, no model) run on every message and every tool call and produce the verdict. When a detector is uncertain, the message is escalated to a review model, which can raise a finding and freeze the agent for the next call but does not sit in the request path. This card covers both.
Intended use
Governing AI agents at the tool boundary: is this call inside the agent's mission, does it move data it should not, does it carry injected instructions. Not intended as a content-moderation system, a general classifier of user text, or a replacement for the model provider's own safety layers.
The review model
Three tiers of Anthropic's Claude models, routed by consequence: a small model for a single soft signal with nothing live, a medium one when the verdict is live or the signals are ambiguous, and the deepest when the verdict can stop an agent or the question is hijack. A tier that declines to assess falls to the next one down and the finding names which reviewer decided. Spend is capped per tenant, per tier and per calling address, and the review is deduplicated per agent and signature for two minutes.
What it sees: the message, the agent's declared mission, and up to 4,000 characters of recent conversation, all redacted the same way stored excerpts are. What it returns: allow, hold or block, with a written rationale that is stored for 30 days and never printed to logs.
Measured accuracy
These are the numbers we have, stated with their limits. The review model was measured once, on 48 labelled uncertain cases from one labeller, on 2026-09-15. The deterministic layer is measured on public corpora in continuous integration with ratchets that fail the build if a number drops.
| Measure | Result | Read it as |
|---|---|---|
| Review model, accuracy at the routed tier | 76% | on the 48 cases it answered; one labeller, so a small and soft number |
| Review model, allow precision / block recall / block precision | 94% / 90% / 64% | it rarely lets an attack through, and about one in three blocks would better have been a hold |
| Review model, latency | 2.9 s median, 7 s p95 | why it is not in the request path |
| Deterministic layer, AgentDojo injections blocked | 27.0% | false positives 2.5% |
| Deterministic layer, InjecAgent injections blocked | 50.8% | 0 of 17 benign flagged |
| Deterministic layer, public jailbreak and injection sets reaching a decision | 25.5% to 50.5% | deepset, jackhhao; false positives 11.8% and 2.0% |
| Deterministic layer, benign instruction corpus (dolly) | 0.0% false positives | on 1,491 ordinary messages |
Limits we state
- —Text detection has a low ceiling against an attacker who adapts. The structural controls (mission, allow-lists, binding, caps, holds) are what bound the damage; the detectors are the cheap first filter and the router to review.
- —The review model is asynchronous. It cannot take back a call the agent already made; it freezes the agent for the next one.
- —The review model is itself a target for injection. Its input is fenced and redacted, and it can only escalate, never widen an agent's scope.
- —Non-English content is escalated, not blocked, unless the fleet declares it expects English.
- —Adaptive attacks against our own stack have not been measured yet. It is the next item on the measurement list, and it needs a model key we run deliberately, not in CI.
Human oversight
With challenge mode on, an uncertain action is held for a person in the dashboard instead of refused; one approval lets exactly one retry through. An agent that parks more than twenty distinct actions in fifteen minutes has further ones refused. Every decision a person makes is stored as a label and feeds the per-detector precision you see on your own fleet.
Changes
Detector changes must keep every ratcheted number or the build fails. Model tier changes are recorded in the release notes. This card is updated when either changes.
Benchmarks
Two public benchmarks with published defense rows. Our protocol removes the model from the loop: a scripted victim executes every injection it reads, so the defense is measured alone. Utility is task completion; attacker success is the injected goal actually executed. Rows from the papers use a real model and cannot be ranked against ours; the two open classifiers appear in both to show how much of a published number was the model.
| Compliant victim, no model | AgentDojo, attacker goals reached | AgentDyn, attacker goals reached | Benign results wrongly withheld |
|---|---|---|---|
| No defense | 597 of 609 | 200 of 200 | 0% |
| ProtectAI classifier | 213 of 609 | 74 of 200 | 30% / 53% |
| PIGuard classifier | — | 100 of 200 | 18% |
| Watchdog, structural layer only | 62 of 609 | 40 of 200 | 0% |
| Watchdog, full | 0 of 609 | 0 of 200 | 3% / 2 to 5% |
The 62 and 40 are hotel bookings and calendar entries, actions that carry nothing to bind; only the text layer catches those. With a current model in the loop the published attacks no longer land even undefended, so a with-model row measures our cost, not our protection; that row is re-run on every detector change and published with the harness.
Vulnerability disclosure policy
We want to hear about security problems in our services, and we will not pursue anyone who finds one while following this policy. This policy is our authorization for that research.
Scope
- —aloviaai.com and its dashboard, API and login.
- —The Watchdog runtime: the check endpoint, the A2A and MCP gateways, the OpenTelemetry ingest, and the Python SDK.
- —Shield: the edge servers and the challenge, block and help pages they serve, when tested against a site you control.
Out of scope
- —Denial of service, rate-limit exhaustion and anything that degrades the service for other customers. Our limits are published and documented; hitting them is not a finding.
- —Social engineering of our staff or customers, phishing, physical attacks.
- —Findings on a customer's own origin, agent or MCP server. Report those to that customer.
- —Third-party services we use (Railway, DigitalOcean, Anthropic, Resend, PostHog). Report those to them.
- —Reports from automated scanners without a demonstrated impact, missing best-practice headers with no exploit, and self-XSS.
Rules
- 01Test only against accounts, fleets and sites you own. Never against another customer's.
- 02Stop as soon as you can demonstrate the issue. Do not read, modify or keep data that is not yours; if you encounter another customer's data, stop and tell us.
- 03No automated scanning at volume, no attempts to disrupt availability, no spam.
- 04Give us a reasonable time to fix before any public disclosure. We ask for 90 days from your report, and we will tell you if we need longer and why.
Safe harbor
Security research conducted in good faith under this policy is authorized. We will not initiate or support legal action, including under the Computer Fraud and Abuse Act or the DMCA, against researchers who comply with it, and we consider such research to be conduct we welcome. If a third party brings action against you for research that complied with this policy, we will make that known. This safe harbor does not cover research that breaks the rules above, and it cannot bind third parties.
How to report
Email security@aloviaai.com. Include the affected surface, steps to reproduce, the impact as you understand it, and how to reach you. We do not publish a PGP key yet; ask in your first message and we will set up an encrypted channel. Reports are also referenced from /.well-known/security.txt.
What to expect
- —Acknowledgement within three business days.
- —Triage and a severity within ten business days, with a fix target: critical and high issues within 30 days, medium within 60, low within 90.
- —Credit on this page if you want it, once the fix is out. There is no paid bounty program yet.
Compliance and legal
- SOC 2 Type II
- Not yet certified. The controls this page describes are the ones an audit would examine; the audit itself has not started.
- ISO/IEC 27001
- Not yet certified.
- Penetration testing
- Internal, owner-run testing and two rounds of independent adversarial review of the Watchdog gateways. No third-party pentest letter yet; one is planned before general availability.
- Breach notice
- Without undue delay to affected customers, with what happened, what data was involved and the steps taken, consistent with California law. See the DPA.
- Privacy
- Privacy policy, including CCPA and CPRA rights and a 45-day response to access and deletion requests.
- Data processing
- Data Processing Addendum and the current subprocessor list, all in the United States.
- Payments
- We do not store cardholder data.
Contact
Security: security@aloviaai.com. Privacy: privacy@aloviaai.com. Everything else: founder@aloviaai.com. Alovia is a service operated by its founder, based in San Francisco, California, United States; no incorporated entity is claimed, and a street address is available on request.
This page as plain text: /trust/trust.txt.