Skip to main content
Version: 2026.09

AI Chat Setup

Istari AI Chat is an assistant inside the web app that answers questions about the resources, systems, and jobs a user can see. It is part of the istari-platform Helm chart and is off until an administrator configures a model. This page covers what an IT administrator must deploy and configure before organization administrators and users can turn it on.

Audience: IT administrators and platform operators. For the end-user feature, see Istari AI Chat. This page reflects the September 2026 release.

Your LLM, your responsibility

AI Chat sends user prompts and platform data to the LLM provider you configure. Review the LLM integration disclaimer and your provider's data-retention terms before enabling it for users.


How it works​

  • The registry service (istari-fileservice) calls the LLM through LiteLLM. The provider is chosen by the prefix of the model string, for example anthropic/…, openai/…, azure/…, or bedrock/….
  • To answer questions about live platform data, the registry service calls the Istari MCP Service (istari-mcp) with the requesting user's own token. Without MCP, the assistant can only chat; it cannot look anything up.
  • The frontend shows the chat entry points (header button and sidebar item) only when the registry service reports that a model is configured. Each user turns them on under Application → Experimental Features; see Istari AI Chat.
  • Users can also enter a personal LLM key in the chat panel (Set an LLM key). It stays in the browser session and travels with each message to the registry service, which forwards it to the provider; the platform never stores it. The key replaces the administrator's credential for that request only; the model and provider endpoint stay as IT configured them, so a user cannot redirect traffic to another provider.

There are two ways to tell the registry service which model to use:

ModeWho configures itChange takes effectUse when
Environment variablesIT administrator, in the registry service secret or Helm valuesAfter a registry service restartOne provider and model for the whole deployment; simplest to stand up
AI Models catalogIT provisions host/credential slots; organization administrators pick hosts and add models in the web appOn the next chat messageSeveral models or gateways, credentials kept out of the admin UI, model changes without a redeploy

Both can coexist. When at least one catalog model is enabled, the catalog is used; otherwise the environment-configured model is used. If neither is set, the API returns 404 and the chat entry points stay hidden. This precedence applies to the streaming endpoint the web app uses, POST /api/v3/ai/chat/stream. The blocking POST /api/v3/ai/chat endpoint, which SDK clients such as the Istari Python client call, reads only the environment variables and ignores the catalog, so SDK integrations need FILE_SERVICE_AI__MODEL even when the catalog is in use.


Prerequisites​

  • Istari Platform installed per Istari Platform Installation, September 2026 release or later.
  • MCP Service enabled and reachable at https://mcp.<customer_istari_fqdn>/mcp/. Follow Optional: MCP Secret and Scenario 1: MCP Service enabled. Required for the assistant to read platform data.
  • An LLM provider you are licensed to use, and either an API key or (for AWS Bedrock) an IAM grant to the registry service pods. See Provider setup.
  • Network egress from the registry service pods to the provider hostname on port 443, and to your own mcp.<customer_istari_fqdn>. See Network requirements.
  • An organization administrator account if you use the AI Models catalog. See Granting admin access.

Option A: Configure with environment variables​

Set FILE_SERVICE_* variables on the registry service. Add them to the istari-fileservice secret you created in Registry Service Secret, or to fileservice.env in istari-values.yaml as in Scenario 7: Custom Environment Variables. Keep API keys in the secret, not in values files.

The only required variable is FILE_SERVICE_AI__MODEL. Everything else depends on the provider.

istari-fileservice-secret.yaml (excerpt)
stringData:
# Required: provider-qualified LiteLLM model string
FILE_SERVICE_AI__MODEL: "anthropic/claude-opus-4-8"
# Provider-dependent (see the tables below)
FILE_SERVICE_AI__API_KEY: "<provider_api_key>"
# Strongly recommended: lets the assistant read platform data
FILE_SERVICE_AI__MCP_SERVER_URL: "https://mcp.<customer_istari_fqdn>/mcp/"

After changing environment variables, restart the registry service:

kubectl rollout restart deployment/istari-fileservice
kubectl rollout status deployment/istari-fileservice

Provider setup​

Anthropic (direct API)​

VariableValue
FILE_SERVICE_AI__MODELanthropic/<model-id>, e.g. anthropic/claude-opus-4-8
FILE_SERVICE_AI__API_KEYAnthropic API key (starts with sk-ant-)

Egress: api.anthropic.com:443. Leave FILE_SERVICE_AI__BASE_URL unset.

OpenAI (direct API)​

VariableValue
FILE_SERVICE_AI__MODELopenai/<model-id>, e.g. openai/gpt-4o
FILE_SERVICE_AI__API_KEYOpenAI API key (starts with sk-)

Egress: api.openai.com:443. Leave FILE_SERVICE_AI__BASE_URL unset.

Azure OpenAI​

VariableValue
FILE_SERVICE_AI__MODELazure/<deployment-name> — the deployment name, not the model name
FILE_SERVICE_AI__BASE_URLhttps://<resource>.openai.azure.com
FILE_SERVICE_AI__API_VERSIONe.g. 2024-02-01
FILE_SERVICE_AI__API_KEYAzure OpenAI key for the resource

Egress: <resource>.openai.azure.com:443.

AWS Bedrock​

The registry service authenticates to Bedrock in one of two ways. Choose one:

  • IAM identity of the pod (recommended). Grant the registry service pod role the permissions in Bedrock IAM permissions via EKS Pod Identity or IRSA. Do not set FILE_SERVICE_AI__API_KEY or FILE_SERVICE_AI__BASE_URL; if set, they bypass the IAM path.
  • Bedrock API key. Set FILE_SERVICE_AI__API_KEY to a long-term Bedrock API key; it is sent as a bearer token.
VariableValue
FILE_SERVICE_AI__MODELbedrock/<model-or-inference-profile-id>, e.g. bedrock/us.anthropic.claude-opus-4-8
FILE_SERVICE_AI__EXTRA_PARAMS{"aws_region_name": "<region>"} — the region whose bedrock-runtime endpoint to call
FILE_SERVICE_AI__API_KEYOnly for the Bedrock API key option

Egress: bedrock-runtime.<region>.amazonaws.com:443. In GovCloud the hostname is bedrock-runtime.us-gov-<west|east>-1.amazonaws.com and the ARN partition is aws-us-gov.

Bedrock also offers an OpenAI-compatible endpoint (https://bedrock-mantle.<region>.api.aws/v1) in some commercial regions. To use it, set FILE_SERVICE_AI__PROVIDER: "standard", FILE_SERVICE_AI__MODEL to the bare model id offered there, FILE_SERVICE_AI__BASE_URL to that URL, and FILE_SERVICE_AI__API_KEY to a bearer token minted for an IAM user with the AmazonBedrockMantleInferenceAccess managed policy. Check AWS documentation for regional availability.

OpenAI-compatible gateway (self-hosted or enterprise)​

For an internal gateway that speaks the OpenAI chat API (vLLM, an enterprise LLM proxy, GenAI.mil, and similar):

VariableValue
FILE_SERVICE_AI__MODELopenai/<model-id-as-the-gateway-names-it>
FILE_SERVICE_AI__BASE_URLhttps://<gateway-host>/v1
FILE_SERVICE_AI__API_KEYGateway credential, if it requires one

If the gateway drops system messages, set FILE_SERVICE_AI__SYSTEM_PROMPT_DELIVERY: "user_turn". If it does not forward the tools API, set FILE_SERVICE_AI__TOOL_MODE: "prompt". Leave both at their defaults for ordinary providers.


Option B: Configure the AI Models catalog​

The catalog splits responsibilities: IT provisions numbered host/credential slots as environment variables; organization administrators choose a host and add models in the web app at Admin → AI Models (/admin/ai-models). Credentials never appear in the UI, and model changes apply on the next chat message without a restart.

IT: provision host slots​

Each slot NN (two digits, 01–99) is a pair of environment variables on the registry service:

VariableRequiredValue
FILE_SERVICE_AI_NN_HOSTYesThe provider hostname, lowercase, without scheme, port, or path (e.g. api.anthropic.com, bedrock-runtime.us-east-1.amazonaws.com); see the notes below for the one accepted URL form
FILE_SERVICE_AI_NN_AUTHNYesThe credential for that host. Set it to an empty string for keyless IAM access (Bedrock). A missing variable makes the slot unusable
FILE_SERVICE_AI_NN_INFONoA short label shown to administrators, useful when two slots share a hostname
istari-fileservice-secret.yaml (excerpt)
stringData:
FILE_SERVICE_AI_01_HOST: "api.anthropic.com"
FILE_SERVICE_AI_01_AUTHN: "<anthropic_api_key>"
FILE_SERVICE_AI_01_INFO: "Anthropic (production key)"
FILE_SERVICE_AI_02_HOST: "bedrock-runtime.us-east-1.amazonaws.com"
FILE_SERVICE_AI_02_AUTHN: "" # empty = use the pod's IAM identity
AWS_REGION: "us-east-1" # Bedrock in Option B: the region LiteLLM calls and signs for
FILE_SERVICE_AI__MCP_SERVER_URL: "https://mcp.<customer_istari_fqdn>/mcp/"

Notes:

  • Set _HOST to the bare hostname. Only https on port 443 is supported; IP addresses, http://, ports, userinfo, and query strings are rejected.
  • The one tolerated exception is a pasted OpenAI-compatible base URL of the form https://<host>/<path>, for example https://gateway.example.com/v1. The registry service uses the hostname and pre-fills /v1 as the default Path for models on that endpoint. Any other extra content makes the slot invalid.
  • The URL path a gateway needs (usually /v1) is configured per model in the Path field, not in _HOST.
  • Slot variables are read at startup. Restart the registry service after adding or changing them.
  • If a slot's hostname later changes, endpoints bound to it show Host changed and fail closed until an administrator re-saves the endpoint.

Organization administrator: add an endpoint and a model​

Organization administrators finish the setup in the web app at Admin → AI Models (/admin/ai-models). The screen-by-screen guide is AI Models (Admin Guide). The points that depend on what IT provisioned:

  • The Host picker lists the slots from the registry service environment. If it is empty, no FILE_SERVICE_AI_NN_HOST / _AUTHN pairs are set, or the pods have not been restarted since they were added.
  • Model ID must be provider-qualified, for example anthropic/claude-opus-4-8, openai/gpt-4o, bedrock/us.anthropic.claude-opus-4-8, or azure/<deployment>.
  • Path is appended to the host: usually /v1 for OpenAI-compatible gateways, blank for Anthropic, Bedrock, and OpenAI. A path pasted into the _HOST value shows up here as the default.
  • Bedrock region: the catalog has no per-model region field in this release. LiteLLM takes the region from the pod's AWS_REGION (or AWS_REGION_NAME) environment variable, so set it explicitly on the registry service, in the secret or in fileservice.env. Do not rely on EKS Pod Identity or IRSA to inject it. If it is missing, LiteLLM falls back to its own default region and calls fail or go to the wrong region.
  • Mark one enabled model as Default; new chats start on it.

Changes apply on the next message from any user. No restart is needed. Prompt text is managed separately under AI Steering (Admin Guide).


Frontend settings​

AI chat is an experimental feature (ai-chat) that each user turns on under Application → Experimental Features. It is available on every deployment; you control its default in the istari-frontend secret (see Frontend Service Secret) or frontend.env:

VariableEffect
VITE_EXPERIMENTAL_DEFAULT_ONComma-separated feature ids that start on for users who have never touched the toggle. Add ai-chat to turn AI chat on by default.
VITE_EXPERIMENTAL_LOCKEDFeature ids pinned to the environment default; the toggle is disabled. Combine with DEFAULT_ON to force on, or use alone to force off.

Restart the frontend deployment after changing these:

kubectl rollout restart deployment/istari-frontend
kubectl rollout status deployment/istari-frontend

If the registry service reports that no model is configured, the chat entry points hide regardless of these settings.


Environment variable reference​

Registry service (istari-fileservice):

VariableRequiredDefaultDescription
FILE_SERVICE_AI__MODELYes (Option A)unsetProvider-qualified LiteLLM model string. Unset with no catalog model = AI Chat disabled
FILE_SERVICE_AI__API_KEYProvider-dependentunsetProvider API key. Not used for Bedrock IAM access
FILE_SERVICE_AI__BASE_URLProvider-dependentunsetEndpoint override. Required for Azure OpenAI and OpenAI-compatible gateways; leave unset for Anthropic, OpenAI, Bedrock
FILE_SERVICE_AI__API_VERSIONAzure only2024-02-01Azure OpenAI API version
FILE_SERVICE_AI__EXTRA_PARAMSBedrock{}JSON object of provider parameters, e.g. {"aws_region_name": "us-east-1"}
FILE_SERVICE_AI__MCP_SERVER_URLStrongly recommendedunsethttps://mcp.<customer_istari_fqdn>/mcp/. Without it the assistant cannot read platform data
FILE_SERVICE_AI__PROVIDERLegacyunsetstandard, anthropic, or azure_openai; only needed with a bare (unprefixed) model name
FILE_SERVICE_AI__SYSTEM_PROMPT_DELIVERYNosystem_roleuser_turn for gateways that drop system messages
FILE_SERVICE_AI__TOOL_MODENonativeprompt for gateways that do not forward the tools API
FILE_SERVICE_AI__MAX_ITERATIONSNo10Maximum tool-call steps per reply before the assistant answers with what it has
FILE_SERVICE_AI__MAX_IDENTICAL_TOOL_CALLSNo3Identical tool calls tolerated before a loop is declared
FILE_SERVICE_AI__DEBUGNofalseVerbose agent logging. Turn on only while troubleshooting
FILE_SERVICE_AI_NN_HOST / _AUTHN / _INFOOption BunsetCatalog host slots; see IT: provision host slots

Variables read by libraries inside the registry service (LiteLLM, the AWS SDK, the HTTP clients), not by the registry service itself:

VariableRead byNotes
HTTPS_PROXY, NO_PROXY, AIOHTTP_TRUST_ENVHTTP clients / LiteLLMSee Proxies
SSL_CERT_FILE, AWS_CA_BUNDLELiteLLM / AWS SDKSee TLS inspection and private CAs
AWS_REGION_NAME, AWS_REGIONLiteLLMBedrock region when aws_region_name is not in FILE_SERVICE_AI__EXTRA_PARAMS. Prefer EXTRA_PARAMS in Option A; in Option B this is the only way to set the region, so set it explicitly (pod identity does not guarantee it)
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKENAWS SDKStatic credentials for Bedrock. Prefer pod identity
AWS_BEARER_TOKEN_BEDROCKLiteLLMFallback Bedrock API key when FILE_SERVICE_AI__API_KEY is unset. Prefer FILE_SERVICE_AI__API_KEY

Frontend service (istari-frontend): VITE_EXPERIMENTAL_DEFAULT_ON, VITE_EXPERIMENTAL_LOCKED — see Frontend settings.


Network requirements​

Egress from the registry service pods​

DestinationPortPurpose
Your LLM provider hostname (see Provider setup)443Chat completions (long-lived streaming responses)
mcp.<customer_istari_fqdn>443Tool calls to the MCP Service. This is your own FQDN, so in-cluster traffic must be able to reach it (hairpin through the load balancer, or split DNS)

The registry service ships with LiteLLM's model catalog bundled and does not download it at startup, so no egress to GitHub or other LiteLLM hosts is needed.

The MCP Service in turn needs its own egress, which depends on whether the API Gateway is enabled:

  • API Gateway enabled (apiGateway.enabled: true): MCP calls api.<customer_istari_fqdn> on port 443 at the /registry and /identity prefixes.
  • No API Gateway: MCP calls registry.<customer_istari_fqdn> and zitadel.<customer_istari_fqdn> directly on port 443.

In both cases the MCP Service itself stays at mcp.<customer_istari_fqdn>; the gateway has no /mcp prefix. These routes already exist if MCP works for external MCP clients.

Firewalls, WAFs, and SNI allow-lists​

  • Allow the provider hostname by SNI/hostname, not by IP; provider IPs change.
  • Chat replies stream over Server-Sent Events from POST /api/v3/ai/chat/stream on the registry service (/registry/api/v3/ai/chat/stream behind the API Gateway). Anything between the browser and the registry (load balancer, WAF, reverse proxy) must not buffer responses and needs an idle timeout of several minutes.
  • The registry service disables HTTP redirects on LLM calls. A proxy that answers with a 30x redirect breaks the call rather than being followed.

Proxies​

The registry service honors the standard HTTPS_PROXY and NO_PROXY variables for most providers. LiteLLM's default transport (used for Bedrock and a few other providers) additionally requires AIOHTTP_TRUST_ENV: "true". Put the cluster-internal hostnames the registry service talks to (PostgreSQL, SpiceDB, NATS, and mcp.<customer_istari_fqdn> if it resolves internally) in NO_PROXY.

note

Validate proxy settings in a staging environment first; proxy behavior differs by product and by provider adapter.

TLS inspection and private CAs​

If outbound traffic passes through TLS inspection, or your provider gateway presents a certificate from a private CA, the registry service must trust that CA or every LLM call fails with CERTIFICATE_VERIFY_FAILED.

  • Preferred: add the CA to the Helm chart's trustedCertBundle as in Scenario 2: Self-Signed TLS Certs.
  • If AI calls still fail, point both SSL_CERT_FILE (used by LiteLLM) and AWS_CA_BUNDLE (used by the AWS SDK) at one PEM bundle that contains the public roots and your inspection CA. A bundle that contains only your private CA breaks object-store and other public TLS calls from the same pod.

Validate the deployment​

Run these from a workstation with kubectl access to the platform namespace. Replace <...> placeholders.

1. The registry service sees the configuration

kubectl exec deployment/istari-fileservice -- sh -c 'env | grep -E "^FILE_SERVICE_AI" | sed "s/=.*/=<set>/"'

FILE_SERVICE_AI__MODEL (or at least one FILE_SERVICE_AI_NN_HOST) and FILE_SERVICE_AI__MCP_SERVER_URL should be listed.

2. Outbound connectivity and TLS from the pod

Probe the provider you configured. Any HTTP status means the pod reached it over TLS; 000 means blocked egress or a TLS failure.

ProviderTest URLExpected
Anthropichttps://api.anthropic.com/v1/models401
OpenAIhttps://api.openai.com/v1/models401
Azure OpenAIhttps://<resource>.openai.azure.com/openai/models?api-version=2024-02-01401
AWS Bedrockhttps://bedrock-runtime.<region>.amazonaws.com/403 or 404
OpenAI-compatible gatewayhttps://<gateway-host>/v1/models401, or 200 if it allows anonymous listing
kubectl exec deployment/istari-fileservice -- curl -sS -o /dev/null -w '%{http_code}\n' "<test-url>"

3. Check for TLS inspection

kubectl exec deployment/istari-fileservice -- curl -sv -o /dev/null "https://<provider-host>" 2>&1 | grep -E 'issuer|SSL certificate'

The issuer should be a public CA. A corporate CA name means traffic is inspected; see TLS inspection and private CAs. SSL certificate problem means the inspecting CA is not trusted yet.

4. The MCP Service is reachable from the pod

kubectl exec deployment/istari-fileservice -- curl -sS -o /dev/null -w '%{http_code}\n' https://mcp.<customer_istari_fqdn>/mcp/

Expected: 401 (reachable; the request carried no token). 404 means the request reached your ingress but there is no route to the MCP Service at /mcp/. A 5xx means the MCP Service is reachable but failing; check the istari-mcp pod logs. 000 means the registry service pod cannot resolve or connect to the MCP FQDN.

5. The AI Chat endpoint answers

In the web app: turn on Istari AI Chat under Application → Experimental Features, click AI chat in the header, and send Reply with the single word OK. A reply confirms the provider path. Then ask What is the name of the system I am looking at? from a system page; a real answer confirms the MCP path.

The command-line test below calls the blocking POST /api/v3/ai/chat endpoint, which reads only the environment variables (Option A). On an Option B-only deployment it returns 404 even though chat works in the web app; use the browser test above instead.

From the command line, you need a bearer token for a platform user. With the Identity Service enabled the platform issues Keys rather than pasteable tokens, so let the Istari Python Client mint the token from a Key you generate under avatar → Developer → Generate Key (see Settings):

export ISTARI_DIGITAL_API_URL="<API URL shown under Developer → Endpoints>"
export ISTARI_CLIENT_IDENTITY_SERVICE_SECRET_FILE="/path/to/downloaded-key.json"
export ISTARI_DIGITAL_IDENTITY_SERVICE_ENABLED=true
AUTH=$(python -c 'from istari_digital_client import Configuration; print(next(v["value"] for v in Configuration().auth_settings().values()))')

curl -sS -X POST "<registry-base-url>/api/v3/ai/chat" \
-H "Authorization: $AUTH" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Reply with the single word OK."}]}'

<registry-base-url> is https://registry.<customer_istari_fqdn> for a direct deployment, or https://api.<customer_istari_fqdn>/registry behind the API Gateway. Where a Personal Access Token still exists, -H "Authorization: Bearer <pat>" works too.

  • 200 with a message object — configured and working.
  • 404 with The AI integration is not configured… — FILE_SERVICE_AI__MODEL is unset (see step 1); expected on an Option B-only deployment.
  • 5xx — the provider call failed; read the registry logs (step 6).

6. Registry service logs

kubectl logs deployment/istari-fileservice --since=15m | grep -i -E 'litellm|ai_service|mcp|CERTIFICATE'

Most provider-side failures reach the user as a generic "Istari AI hit an error" message; the specific cause is only in these logs.


Troubleshooting​

SymptomLikely causeWhat to do
Toggle is on but the AI chat button is missing, or a toast says Istari AI is not configured on this serverNo FILE_SERVICE_AI__MODEL and no enabled catalog model; or env vars set but the pod not restartedSet the model (Option A) or enable a model (Option B); kubectl rollout restart deployment/istari-fileservice after env changes
Assistant chats but cannot see systems or files; replies contain text that looks like tool callsFILE_SERVICE_AI__MCP_SERVER_URL unset, wrong path (must end in /mcp/), MCP not deployed, or the registry cannot reach mcp.<fqdn>Set the URL, run validation step 4, check istari-mcp pod logs; fix hairpin/split DNS so pods reach your own FQDN
Registry logs show CERTIFICATE_VERIFY_FAILED / unable to get local issuer certificateTLS inspection or a private CA on the provider pathRun validation step 3; add the CA via trustedCertBundle, or set SSL_CERT_FILE and AWS_CA_BUNDLE to a full bundle
Connection timeouts or 000 from validation step 2Egress blocked (security group, NAT, firewall, WAF SNI list) or proxy not configuredAllow the provider FQDN on 443; set HTTPS_PROXY/NO_PROXY/AIOHTTP_TRUST_ENV if a proxy is mandatory
Provider returns 401/403 (in registry logs)Wrong key for the provider (sk-ant- vs sk-), key for another resource/region, revoked keyRe-check the key and provider match; for Azure confirm the key belongs to the resource in BASE_URL
Bedrock AccessDeniedExceptionPod role lacks bedrock:InvokeModel*, model access not enabled in that region, marketplace subscription missingApply Bedrock IAM permissions; enable model access in the Bedrock console for that region; subscribe to the model in AWS Marketplace
Bedrock ValidationException / model not foundWrong model id or region; cross-region models need the us./eu. inference-profile idUse the inference profile id (e.g. bedrock/us.anthropic.claude-opus-4-8); set aws_region_name (Option A) or the pod's AWS_REGION (Option B) to a region that serves it
Azure 404 / DeploymentNotFoundFILE_SERVICE_AI__MODEL uses the model name instead of the deployment name, or API_VERSION missingUse azure/<deployment-name>; set FILE_SERVICE_AI__API_VERSION
Provider 400 mentioning temperature or an unsupported parameter for a brand-new modelThe platform's bundled LiteLLM predates the modelPick a model the current release supports, or upgrade the platform
AI Models page: Host picker is emptyNo FILE_SERVICE_AI_NN_HOST/_AUTHN pairs on the registry service, or _AUTHN missing (not empty)Provision slots per Option B and restart the registry service
Endpoint badge Host changed / Host not provisioned / Invalid host configSlot hostname changed, slot removed, or _HOST value fails validationRestore or fix the slot value; an administrator re-saves the endpoint to confirm
Replies stop mid-sentence; "The connection dropped before the reply finished"A proxy or load balancer buffers or times out the SSE streamDisable response buffering and raise idle timeouts on the path to /api/v3/ai/chat/stream
Provider 429Provider rate limit or quotaRaise the quota with the provider or choose a different model

Bedrock IAM permissions​

Under review

This section is being finalized with Istari Digital's infrastructure team. Validate it against your account before relying on it.

To call Bedrock with the pod's IAM identity (no API key):

1. Enable model access for the models you want in the Bedrock console for the region you set in aws_region_name. Third-party models (Anthropic, Meta, and others) also require an active AWS Marketplace subscription in the account.

2. Attach a policy to the IAM role the registry service pods assume (EKS Pod Identity or IRSA). The template below covers one configured model id, <model-id>, for example anthropic.claude-opus-4-8 or amazon.nova-pro-v1:0. Replace <partition> (aws, or aws-us-gov in GovCloud), <region> (the region you call), <account-id>, <geo>, and the destination regions, then remove the resources that do not apply to how you invoke the model.

{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "BedrockInvokeModel",
"Effect": "Allow",
"Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
"Resource": [
"arn:<partition>:bedrock:<region>::foundation-model/<model-id>",
"arn:<partition>:bedrock:<region>:<account-id>:inference-profile/<geo>.<model-id>",
"arn:<partition>:bedrock:<destination-region-1>::foundation-model/<model-id>",
"arn:<partition>:bedrock:<destination-region-2>::foundation-model/<model-id>"
]
},
{
"Sid": "BedrockMarketplaceViewSubscriptions",
"Effect": "Allow",
"Action": ["aws-marketplace:ViewSubscriptions"],
"Resource": "*"
}
]
}
  • Direct model id (bedrock/<model-id>, no geographic prefix): only the first resource is needed, the foundation model in the region you call.

  • Cross-region inference profile (bedrock/<geo>.<model-id>, where <geo> is us, eu, apac, and so on): needs the profile ARN in the region you call and a foundation-model ARN for every region the profile routes to. List them with:

    aws bedrock get-inference-profile --region <region> --inference-profile-identifier <geo>.<model-id> --query 'models[].modelArn'
  • The template names one model. Repeat the resources for each model id you configure, or use a prefix wildcard where your policy allows, for example arn:<partition>:bedrock:*::foundation-model/anthropic.claude-* for every region and Claude model.

  • aws-marketplace:ViewSubscriptions lets the invoking principal verify the account's subscription for marketplace-gated models (Anthropic and other third-party models). Subscribing itself is a one-time account action, not something the pod role needs.

3. Optional: keep traffic inside the VPC with interface endpoints for bedrock-runtime (and bedrock) with private DNS enabled. No registry service change is needed; the same hostname resolves to the endpoint.

If your platform cannot use pod identity, create an IAM user with the same policy and either generate a Bedrock API key for it (set as FILE_SERVICE_AI__API_KEY, or as a slot's _AUTHN) or use the OpenAI-compatible endpoint described under AWS Bedrock.