AI Reserve Developer Documentation

Trust & Disclosures

Data Privacy & Security Overview

How we protect client data across the AI Reserve platform. Every client's data is treated as confidential by default: it is logically isolated per client, encrypted in transit and at rest, and never used to train shared AI models. This page explains the principles behind that design and the controls that enforce them, at the two layers a security team will want to evaluate — AI Reserve's own platform (the gateway and supporting services we build and operate on Google Cloud) and the upstream model providers that run inference on your prompts under their own terms.

Last reviewed: September 28, 2026

What AI Reserve is


AI Reserve is a gateway between your applications and AI model providers: applications integrate once and route requests to the models your organization approves. Because the gateway speaks the OpenAI, Anthropic, and Bedrock wire formats, it also works as a drop-in backend for common AI developer tools such as Cursor — so your team's coding-assistant traffic can run through the same per-tenant allowlist, routing, and zero-retention controls as your applications.

For the purposes of this overview, the focus is the data path — the gateway is a thin routing layer, not a data store. API prompts and responses are not persisted (how AI models are used), and each organization's traffic can reach only the providers on its own allowlist, enforced in our model router (client data stays separated).

Media generation follows the same rule. Large media inputs for image- and video-generation models (source images, reference audio or video) are staged in client-namespaced transient storage and become eligible for automatic deletion at 24 hours via a storage lifecycle rule — they are inference inputs, not storage. Inline media payloads are never persisted into stored job records: they are replaced with placeholders before the record is written. Media generation content (prompts, generation parameters, media references) is not retained: job records are scrubbed to billing metadata — model, duration, cost, status, provider job reference, timestamps — when the job completes, with no setting that defers it (see the retention table below).

Architecture


All AI Reserve infrastructure runs in our own Google Cloud project (US region today). The full stack is defined as infrastructure-as-code, which keeps environments reproducible and auditable — and makes a region-pinned EU deployment a scoped exercise if your organization ever requires it.

Usage metadata. Everything AI Reserve stores lives in that Google Cloud project: Cloud SQL (PostgreSQL) for account and usage data — including any organization-wide instructions and context your admins choose to define, short admin-written guidance text applied to your organization's traffic — and Google Cloud Storage for documents uploaded through the platform's optional web chat and document features, a surface API-only organizations don't use.

Data is encrypted at rest and in transit (TLS). Credentials are held in GCP Secret Manager, authentication is handled by Firebase Auth, and administrative and access events are audit-logged (encryption & auditability). The underlying Google Cloud platform carries SOC 2 Type II and ISO 27001 — those are Google's certifications for the infrastructure layer (compliance & certifications).

Where data goes. The only egress of customer content is model inference over TLS to the providers you select. Enrichment sources, where used, receive only public company identifiers — never your prompt content.

The architecture described here reflects the platform as of the date above. AI Reserve may evolve implementation details over time; any changes material to the data-handling commitments on this page will be communicated.

Client data stays separated


Each client operates in its own logical boundary, and nothing crosses that boundary implicitly:

  • Per-tenant isolation. Access to a client's workspace requires explicit authorization scoped to that client. Each AI request is processed independently — there is no shared memory or cross-client visibility at any step, and outputs are returned to the originating client's workspace only.
  • Per-tenant routing policy. Your provider allowlist means no provider outside your approved set ever sees your traffic.
  • Least-privilege staff access. AI Reserve staff do not review client conversations or content in the ordinary course of business. Internal tools and team members reach only the specific client data a task requires; access is not granted by default and access events are audit-logged.

Encryption & auditability — everywhere data moves (or rests)


  • In transit: TLS on all connections — from your applications to the gateway, and from the gateway to every model provider. There is no unencrypted path for client data to travel.
  • At rest: all stored data (Cloud SQL, Cloud Storage) is encrypted at rest on Google Cloud.
  • Secrets: provider credentials and API keys are held in GCP Secret Manager — never in code or config files.
  • Authentication: Firebase Auth for platform users; Google SSO internally for AI Reserve staff.
  • Audit logging: pgAudit on Cloud SQL. Audit logs capture administrative and access events — who did what, when — never prompt or response content.
  • Engineering supply chain: the GitHub organization is hardened with enforced two-factor authentication (organization-level enforcement), automated dependency updates, CodeQL static analysis, and secret scanning with push protection.

How AI models are used


AI processing is a pass-through step, not a place where client data accumulates or is learned from.

At the gateway: zero retention of API content

For API traffic AI Reserve keeps zero retention of content, with no exception: prompts and responses pass through the gateway in memory only and are never written to disk, and no setting or support procedure can change that. What we store is usage metadata — model, token counts, cost — for billing and analytics, never the content itself. Client data is never used to train shared AI models. Asynchronous media generation is the one place a request is held while work is in progress: the job record keeps the prompt and generation settings only while the job is running and is scrubbed to billing metadata the moment it completes, as described above and in the retention table. (The platform's optional web chat surface stores conversation content by default to power chat history — an organization can turn that storage off entirely; see the retention table.)

An optional usage-classification feature categorizes AI spend by business function by classifying the standing purpose of each API key and team from its name, description, and organizational context (details your organization sets itself, like team names, structure, and which user or team a key belongs to) — by default it does not read message content — and it is off for newly created organizations until an administrator explicitly enables it. A separate per-request mode, off unless an administrator turns it on, classifies each request from a bounded excerpt of its prompt that is processed in memory and never stored (see the data-handling page).

At the model providers: no training, bounded retention

When your traffic reaches a model provider, that provider's terms govern. Every provider reachable through default routing is documented as not training on API inputs or outputs under the terms we route under, with disclosed, bounded retention (typical defaults are on the order of 30 days; several providers retain nothing by default). The catalog spans frontier models from every major provider — served directly by the model makers or via vetted US GPU clouds and aggregators — and the routing configuration is checked against these documented postures by deterministic checks that run in CI on every merge, plus a scheduled weekday registry sync.

The per-provider matrix — training posture, retention window, processing region, and DPA role for every provider we route to — is maintained on the Provider Data Handling page, which is the authoritative source for those characteristics. If a provider fails or is saturated, requests may be retried only through a disclosed failover chain of identical-model-weights deployments on disclosed subprocessors, and every fallback serve is disclosed per response — see provider failover on that page.

Providers that soften a platform guarantee are locked by default for every organization and stay locked until an organization administrator explicitly opts in: providers whose terms permit training on inputs (for example DeepSeek's direct API and Moonshot AI) are independent controllers reachable only behind an explicit training consent, and new organizations start with PRC-hosted providers disabled until an administrator opts in. See the consent-gates section of the data-handling page.

Strict zero-data-retention configurations

For organizations that must guarantee zero data retention end to end, we pin the organization to provider-served Claude models — AWS Bedrock or Google Vertex AI today — and block the direct-API variants at the organization level via the provider allowlist. Zero-retention options are also documented for several of the vetted GPU clouds serving open-weight models, and stricter ZDR arrangements with providers such as OpenAI and Anthropic can be arranged per organization through each provider's approval process (model-family eligibility varies). Current per-provider ZDR status lives on the Provider Data Handling page; if your organization requires ZDR on a specific provider, contact us.

Retention and deletion are client-controlled


Data Retention
Prompts & responses (API traffic) Not stored — in memory only during request handling. There is no capture mechanism: no administrator control, support procedure, or setting can turn on storage of API prompts or responses.
Prompts & responses (optional web chat) Stored by default to power chat history and search. An organization can turn content storage off entirely, permanently purge stored content on demand, or set an automatic retention window that deletes stored content on a schedule.
Media-generation content (prompt, generation parameters, media references) Held only while the job is running — the job record is scrubbed to billing metadata (model, duration, cost, status, provider job reference, timestamps) the moment the job completes; no setting defers the scrub. Staged media inputs auto-delete from transient storage (eligible at 24 hours); inline media payloads are never persisted in job records.
Usage metadata (model, token counts, cost) Retained for 24 months after account closure (administrative and access audit logs likewise); invoices and ledger entries for three years — for billing integrity and audit, per Data Processing Agreement §8.3.
Client account data Life of the relationship; removed on offboarding.
Database backups Roll off after 7 days.
Documents saved to your document library (from the documents page, or when you choose "Store for future chats" after attaching one in chat) Retained until the user or organization deletes them. Chat can search saved documents to answer questions — to power that search, each saved document is also indexed in a Google Gemini File Search store scoped to your organization; the platform deletes that indexed copy when you delete the document, a daily sweep removes any copy a deletion could not complete, and the store is deleted with the organization (see data handling, including the clean-up of copies an earlier defect left behind). API traffic never adds to the library.
Files attached in web chat without saving (images, and documents attached with the "Just this conversation" option) Stored with the conversation and removed by the same controls that delete chat content — deleting the chat, "Delete stored content," the automatic deletion window, or deleting the account. Not retained when the organization has content storage turned off. API traffic does not create chat attachments.
Access & audit logs Retained for security review; not client-editable by design.

Deletion follows our Data Processing Agreement terms: customer data is deleted within a 14-day window of a verified request, to the extent technically possible, subject to legal-retention carve-outs. Billing and audit records are retained as set out in the DPA. Full data deletion is available on account offboarding. Deletion requests can be directed to your AI Reserve point of contact or security@aireserve.com.

Compliance and certifications


Infrastructure layer (Google Cloud): SOC 2 Type II and ISO 27001 — Google's certifications covering the platform AI Reserve runs on, not AI Reserve's own attestations.

AI Reserve's own program: our SOC 2 Type II examination is in progress (Type I first, with the Type II observation window to follow) and no report has been issued yet — we do not claim our own certification. The control set is implemented and documented, compliance automation (Vanta) runs continuous checks against the live environment, and our build pipeline enforces our public claims registries on every change. Once a report is issued, our Data Processing Agreement provides that a current report is accepted in lieu of an audit of covered controls.

The full certification matrix — our platform, the Google Cloud infrastructure layer, the software supply chain, and every model serving provider, each certification attributed to the organization that holds it — is on the Security Certifications page; per-provider training, retention, and region posture is on the Provider Data Handling page.

Subprocessors


AI Reserve engages a small set of subprocessors: Google Cloud Platform (hosting, storage, databases), identity providers (Google Firebase Authentication; Auth0 for keyless desktop sign-in), Resend (transactional email — account-related emails only; never receives prompt or response content), Slack (optional Slack Connect customer channels — the messages participants choose to post there, including agent prompts and completions for organizations using the @aireserve mention agent), Microsoft (optional Microsoft Teams customer channels — business communications only, never prompt or response content — and Microsoft Entra ID sign-in for the Microsoft Office integration), Daytona (isolated, single-use sandboxes for the optional Slack code agent — the code request, its thread context, and the approved repository's contents, only for organizations whose administrator installs the AI Reserve GitHub App; listed ahead of first production use, not yet in use for any organization), RunPod (customer-configured GPU pods and volumes for the GPU compute feature — enabled in our production and UAT environments for organizations whose administrator has turned GPU compute on; no organization has it enabled yet, and never gateway prompts or responses), Apollo.io (business-contact and company lookups for the optional chat personas and company profile pages — listed ahead of enablement; not enabled, and the web app, where every deployed Apollo call site lives, holds no credential for it), Exa (web search for the portal chat — the search query the model derives from the user's prompt, and nothing else; never other conversation content, documents, or account data), Composio (optional app connectors — for users who connect their own applications from the Tools page, the authorization for those applications and the connector tool calls and results made on the user's behalf in a conversation the user attached them to; never API traffic), and the model providers your organization selects, which receive prompt content only when your organization's routing policy sends traffic to them. Providers whose terms permit training on inputs are independent controllers, not subprocessors, and are reachable only behind explicit organization consent.

The authoritative, always-current list — with each subprocessor's function and jurisdiction, as referenced in Annex 5 of our Data Processing Addendum — is published on the Subprocessors page.

Our commitments on training, sharing, and deletion


These are the commitments in our Terms of Service and Data Processing Agreement:

  • Training: identifiable customer data is never used by AI Reserve to train models; any product-improvement use is limited to de-identified, aggregated data as set out in the DPA.
  • Sharing: customer data is never sold. The only customer data that leaves the platform is:
    • prompts and responses, sent to the model providers your organization has selected in the performance of the Services;
    • account email addresses, sent to our email delivery provider (a subprocessor used only to send account-related emails; it never receives prompt or response content);
    • hosting data processed by our disclosed infrastructure subprocessors as part of running the platform;
    • media inputs supplied for image, audio, and video generation, sent to the media provider your organization has selected for that generation;
    • where the voice-cloning capability is enabled for your organization, voice recordings and derived voice models, sent to the provider identified in the DPA (Annex 1);
    • request IP addresses, sent to our geolocation and anonymizer screening subprocessor for origin screening on the machine API; and
    • account identifiers and authentication data, sent to our authentication subprocessors, as identified on the Subprocessor List, to authenticate users and control access to the Services.
    Nothing is shared for advertising, analytics, or any other purpose. Disclosure may occur where required by law. For organizations that adopt a collaboration surface (for example, a Slack Connect shared channel, including its optional mention agent), the messages participants choose to post in that surface and the completions returned there are processed by the operating subprocessor as disclosed on the Subprocessor List. Where the mention agent is invoked in a message thread, or responds to a follow-up in a thread in which it is already participating, the messages in that thread (including messages posted by participants who did not invoke the agent) are included as context in the request sent to the selected model provider, on the same terms as other prompt content under the first bullet above. Your API traffic never passes through a collaboration surface; only content participants themselves post there is processed by that subprocessor. For organizations that use the web document library, the content of each saved document is sent to the Google Gemini API — the model provider that categorizes it, extracts searchable text, and indexes it for chat search in a File Search store scoped to your organization — and that indexed copy stays with the provider until the document is deleted, as disclosed on the Subprocessor List and in the saved-documents section of the Provider Data Handling page. For organizations whose administrator installs the AI Reserve GitHub App and approves a repository for the Slack code agent, each explicit code request — the request text, the surrounding Slack thread context, and the approved repository's contents — is processed inside one isolated, single-use sandbox operated by Daytona, our sandbox subprocessor, as disclosed on the Subprocessor List; the sandbox is deleted at the end of the run, and the model inference for that request is served through our gateway by the model provider your organization has selected, on the same terms as other prompt content under the first bullet above. When the portal chat assistant searches the web to answer a request, the search query the model derives from the user's prompt — and nothing else — is sent to Exa, our web search subprocessor, which returns public web results, as disclosed on the Subprocessor List. For users who connect their own applications from the portal's Tools page, Composio, our connector subprocessor, holds the authorization the user grants and, in a conversation the user attaches a connected application to, relays the connector tool calls the assistant makes and the application's responses on the user's behalf, as disclosed on the Subprocessor List; each connection is user-initiated and revocable by the user at any time.
  • Deletion: on a verified request, the customer data stored — account data, uploaded documents, and any opt-in stored content — is deleted within a 14-day window, to the extent technically possible, with legal-retention carve-outs; billing and audit records are retained as set out in the DPA. For API-only organizations this deletion commitment applies chiefly to account data — prompt and response content is never stored to begin with.

What we don't do


  • We don't sell or share client data with third parties for their own purposes.
  • We don't train shared models on any client's proprietary information.
  • We don't allow one client's environment to be visible to another.
  • We don't retain data beyond what's needed once a client asks for deletion, subject only to the legal-retention carve-outs in the DPA.
  • We don't route your traffic to any provider outside your organization's approved allowlist.

Questions about this overview or a security review: contact your AI Reserve point of contact or security@aireserve.com.