AI Chat Setup
Istari AI Chat is an assistant inside the web app that answers questions about the resources, systems, and jobs a user can see. It is part of the istari-platform Helm chart and is off until an administrator configures a model. This page covers what an IT administrator must deploy and configure before organization administrators and users can turn it on.
Audience: IT administrators and platform operators. For the end-user feature, see Istari AI Chat. This page reflects the September 2026 release.
AI Chat sends user prompts and platform data to the LLM provider you configure. Review the LLM integration disclaimer and your provider's data-retention terms before enabling it for users.
How it works
- The registry service (
istari-fileservice) calls the LLM through LiteLLM. The provider is chosen by the prefix of the model string, for exampleanthropic/…,openai/…,azure/…, orbedrock/…. - To answer questions about live platform data, the registry service calls the Istari MCP Service (
istari-mcp) with the requesting user's own token. Without MCP, the assistant can only chat; it cannot look anything up. - The frontend shows the chat entry points (header button and sidebar item) only when the registry service reports that a model is configured. Each user turns them on under Application → Experimental Features; see Istari AI Chat.
- Users can also enter a personal LLM key in the chat panel (Set an LLM key). It stays in the browser session and travels with each message to the registry service, which forwards it to the provider; the platform never stores it. The key replaces the administrator's credential for that request only; the model and provider endpoint stay as IT configured them, so a user cannot redirect traffic to another provider.
There are two ways to tell the registry service which model to use:
| Mode | Who configures it | Change takes effect | Use when |
|---|---|---|---|
| Environment variables | IT administrator, in the registry service secret or Helm values | After a registry service restart | One provider and model for the whole deployment; simplest to stand up |
| AI Models catalog | IT provisions host/credential slots; organization administrators pick hosts and add models in the web app | On the next chat message | Several models or gateways, credentials kept out of the admin UI, model changes without a redeploy |
Both can coexist. When at least one catalog model is enabled, the catalog is used; otherwise the environment-configured model is used. If neither is set, the API returns 404 and the chat entry points stay hidden. This precedence applies to the streaming endpoint the web app uses, POST /api/v3/ai/chat/stream. The blocking POST /api/v3/ai/chat endpoint, which SDK clients such as the Istari Python client call, reads only the environment variables and ignores the catalog, so SDK integrations need FILE_SERVICE_AI__MODEL even when the catalog is in use.
Prerequisites
- Istari Platform installed per Istari Platform Installation, September 2026 release or later.
- MCP Service enabled and reachable at
https://mcp.<customer_istari_fqdn>/mcp/. Follow Optional: MCP Secret and Scenario 1: MCP Service enabled. Required for the assistant to read platform data. - An LLM provider you are licensed to use, and either an API key or (for AWS Bedrock) an IAM grant to the registry service pods. See Provider setup.
- Network egress from the registry service pods to the provider hostname on port 443, and to your own
mcp.<customer_istari_fqdn>. See Network requirements. - An organization administrator account if you use the AI Models catalog. See Granting admin access.
Option A: Configure with environment variables
Set FILE_SERVICE_* variables on the registry service. Add them to the istari-fileservice secret you created in Registry Service Secret, or to fileservice.env in istari-values.yaml as in Scenario 7: Custom Environment Variables. Keep API keys in the secret, not in values files.
The only required variable is FILE_SERVICE_AI__MODEL. Everything else depends on the provider.
stringData:
# Required: provider-qualified LiteLLM model string
FILE_SERVICE_AI__MODEL: "anthropic/claude-opus-4-8"
# Provider-dependent (see the tables below)
FILE_SERVICE_AI__API_KEY: "<provider_api_key>"
# Strongly recommended: lets the assistant read platform data
FILE_SERVICE_AI__MCP_SERVER_URL: "https://mcp.<customer_istari_fqdn>/mcp/"
After changing environment variables, restart the registry service:
kubectl rollout restart deployment/istari-fileservice
kubectl rollout status deployment/istari-fileservice
Provider setup
Anthropic (direct API)
| Variable | Value |
|---|---|
FILE_SERVICE_AI__MODEL | anthropic/<model-id>, e.g. anthropic/claude-opus-4-8 |
FILE_SERVICE_AI__API_KEY | Anthropic API key (starts with sk-ant-) |
Egress: api.anthropic.com:443. Leave FILE_SERVICE_AI__BASE_URL unset.
OpenAI (direct API)
| Variable | Value |
|---|---|
FILE_SERVICE_AI__MODEL | openai/<model-id>, e.g. openai/gpt-4o |
FILE_SERVICE_AI__API_KEY | OpenAI API key (starts with sk-) |
Egress: api.openai.com:443. Leave FILE_SERVICE_AI__BASE_URL unset.
Azure OpenAI
| Variable | Value |
|---|---|
FILE_SERVICE_AI__MODEL | azure/<deployment-name> — the deployment name, not the model name |
FILE_SERVICE_AI__BASE_URL | https://<resource>.openai.azure.com |
FILE_SERVICE_AI__API_VERSION | e.g. 2024-02-01 |
FILE_SERVICE_AI__API_KEY | Azure OpenAI key for the resource |
Egress: <resource>.openai.azure.com:443.
AWS Bedrock
The registry service authenticates to Bedrock in one of two ways. Choose one:
- IAM identity of the pod (recommended). Grant the registry service pod role the permissions in Bedrock IAM permissions via EKS Pod Identity or IRSA. Do not set
FILE_SERVICE_AI__API_KEYorFILE_SERVICE_AI__BASE_URL; if set, they bypass the IAM path. - Bedrock API key. Set
FILE_SERVICE_AI__API_KEYto a long-term Bedrock API key; it is sent as a bearer token.
| Variable | Value |
|---|---|
FILE_SERVICE_AI__MODEL | bedrock/<model-or-inference-profile-id>, e.g. bedrock/us.anthropic.claude-opus-4-8 |
FILE_SERVICE_AI__EXTRA_PARAMS | {"aws_region_name": "<region>"} — the region whose bedrock-runtime endpoint to call |
FILE_SERVICE_AI__API_KEY | Only for the Bedrock API key option |
Egress: bedrock-runtime.<region>.amazonaws.com:443. In GovCloud the hostname is bedrock-runtime.us-gov-<west|east>-1.amazonaws.com and the ARN partition is aws-us-gov.
Bedrock also offers an OpenAI-compatible endpoint (https://bedrock-mantle.<region>.api.aws/v1) in some commercial regions. To use it, set FILE_SERVICE_AI__PROVIDER: "standard", FILE_SERVICE_AI__MODEL to the bare model id offered there, FILE_SERVICE_AI__BASE_URL to that URL, and FILE_SERVICE_AI__API_KEY to a bearer token minted for an IAM user with the AmazonBedrockMantleInferenceAccess managed policy. Check AWS documentation for regional availability.
OpenAI-compatible gateway (self-hosted or enterprise)
For an internal gateway that speaks the OpenAI chat API (vLLM, an enterprise LLM proxy, GenAI.mil, and similar):
| Variable | Value |
|---|---|
FILE_SERVICE_AI__MODEL | openai/<model-id-as-the-gateway-names-it> |
FILE_SERVICE_AI__BASE_URL | https://<gateway-host>/v1 |
FILE_SERVICE_AI__API_KEY | Gateway credential, if it requires one |
If the gateway drops system messages, set FILE_SERVICE_AI__SYSTEM_PROMPT_DELIVERY: "user_turn". If it does not forward the tools API, set FILE_SERVICE_AI__TOOL_MODE: "prompt". Leave both at their defaults for ordinary providers.
Option B: Configure the AI Models catalog
The catalog splits responsibilities: IT provisions numbered host/credential slots as environment variables; organization administrators choose a host and add models in the web app at Admin → AI Models (/admin/ai-models). Credentials never appear in the UI, and model changes apply on the next chat message without a restart.
IT: provision host slots
Each slot NN (two digits, 01–99) is a pair of environment variables on the registry service:
| Variable | Required | Value |
|---|---|---|
FILE_SERVICE_AI_NN_HOST | Yes | The provider hostname, lowercase, without scheme, port, or path (e.g. api.anthropic.com, bedrock-runtime.us-east-1.amazonaws.com); see the notes below for the one accepted URL form |
FILE_SERVICE_AI_NN_AUTHN | Yes | The credential for that host. Set it to an empty string for keyless IAM access (Bedrock). A missing variable makes the slot unusable |
FILE_SERVICE_AI_NN_INFO | No | A short label shown to administrators, useful when two slots share a hostname |
stringData:
FILE_SERVICE_AI_01_HOST: "api.anthropic.com"
FILE_SERVICE_AI_01_AUTHN: "<anthropic_api_key>"
FILE_SERVICE_AI_01_INFO: "Anthropic (production key)"
FILE_SERVICE_AI_02_HOST: "bedrock-runtime.us-east-1.amazonaws.com"
FILE_SERVICE_AI_02_AUTHN: "" # empty = use the pod's IAM identity
AWS_REGION: "us-east-1" # Bedrock in Option B: the region LiteLLM calls and signs for
FILE_SERVICE_AI__MCP_SERVER_URL: "https://mcp.<customer_istari_fqdn>/mcp/"
Notes:
- Set
_HOSTto the bare hostname. Onlyhttpson port 443 is supported; IP addresses,http://, ports, userinfo, and query strings are rejected. - The one tolerated exception is a pasted OpenAI-compatible base URL of the form
https://<host>/<path>, for examplehttps://gateway.example.com/v1. The registry service uses the hostname and pre-fills/v1as the default Path for models on that endpoint. Any other extra content makes the slot invalid. - The URL path a gateway needs (usually
/v1) is configured per model in the Path field, not in_HOST. - Slot variables are read at startup. Restart the registry service after adding or changing them.
- If a slot's hostname later changes, endpoints bound to it show Host changed and fail closed until an administrator re-saves the endpoint.
Organization administrator: add an endpoint and a model
Organization administrators finish the setup in the web app at Admin → AI Models (/admin/ai-models). The screen-by-screen guide is AI Models (Admin Guide). The points that depend on what IT provisioned:
- The Host picker lists the slots from the registry service environment. If it is empty, no
FILE_SERVICE_AI_NN_HOST/_AUTHNpairs are set, or the pods have not been restarted since they were added. - Model ID must be provider-qualified, for example
anthropic/claude-opus-4-8,openai/gpt-4o,bedrock/us.anthropic.claude-opus-4-8, orazure/<deployment>. - Path is appended to the host: usually
/v1for OpenAI-compatible gateways, blank for Anthropic, Bedrock, and OpenAI. A path pasted into the_HOSTvalue shows up here as the default. - Bedrock region: the catalog has no per-model region field in this release. LiteLLM takes the region from the pod's
AWS_REGION(orAWS_REGION_NAME) environment variable, so set it explicitly on the registry service, in the secret or infileservice.env. Do not rely on EKS Pod Identity or IRSA to inject it. If it is missing, LiteLLM falls back to its own default region and calls fail or go to the wrong region. - Mark one enabled model as Default; new chats start on it.
Changes apply on the next message from any user. No restart is needed. Prompt text is managed separately under AI Steering (Admin Guide).
Frontend settings
AI chat is an experimental feature (ai-chat) that each user turns on under Application → Experimental Features. It is available on every deployment; you control its default in the istari-frontend secret (see Frontend Service Secret) or frontend.env:
| Variable | Effect |
|---|---|
VITE_EXPERIMENTAL_DEFAULT_ON | Comma-separated feature ids that start on for users who have never touched the toggle. Add ai-chat to turn AI chat on by default. |
VITE_EXPERIMENTAL_LOCKED | Feature ids pinned to the environment default; the toggle is disabled. Combine with DEFAULT_ON to force on, or use alone to force off. |
Restart the frontend deployment after changing these:
kubectl rollout restart deployment/istari-frontend
kubectl rollout status deployment/istari-frontend
If the registry service reports that no model is configured, the chat entry points hide regardless of these settings.
Environment variable reference
Registry service (istari-fileservice):
| Variable | Required | Default | Description |
|---|---|---|---|
FILE_SERVICE_AI__MODEL | Yes (Option A) | unset | Provider-qualified LiteLLM model string. Unset with no catalog model = AI Chat disabled |
FILE_SERVICE_AI__API_KEY | Provider-dependent | unset | Provider API key. Not used for Bedrock IAM access |
FILE_SERVICE_AI__BASE_URL | Provider-dependent | unset | Endpoint override. Required for Azure OpenAI and OpenAI-compatible gateways; leave unset for Anthropic, OpenAI, Bedrock |
FILE_SERVICE_AI__API_VERSION | Azure only | 2024-02-01 | Azure OpenAI API version |
FILE_SERVICE_AI__EXTRA_PARAMS | Bedrock | {} | JSON object of provider parameters, e.g. {"aws_region_name": "us-east-1"} |
FILE_SERVICE_AI__MCP_SERVER_URL | Strongly recommended | unset | https://mcp.<customer_istari_fqdn>/mcp/. Without it the assistant cannot read platform data |
FILE_SERVICE_AI__PROVIDER | Legacy | unset | standard, anthropic, or azure_openai; only needed with a bare (unprefixed) model name |
FILE_SERVICE_AI__SYSTEM_PROMPT_DELIVERY | No | system_role | user_turn for gateways that drop system messages |
FILE_SERVICE_AI__TOOL_MODE | No | native | prompt for gateways that do not forward the tools API |
FILE_SERVICE_AI__MAX_ITERATIONS | No | 10 | Maximum tool-call steps per reply before the assistant answers with what it has |
FILE_SERVICE_AI__MAX_IDENTICAL_TOOL_CALLS | No | 3 | Identical tool calls tolerated before a loop is declared |
FILE_SERVICE_AI__DEBUG | No | false | Verbose agent logging. Turn on only while troubleshooting |
FILE_SERVICE_AI_NN_HOST / _AUTHN / _INFO | Option B | unset | Catalog host slots; see IT: provision host slots |
Variables read by libraries inside the registry service (LiteLLM, the AWS SDK, the HTTP clients), not by the registry service itself:
| Variable | Read by | Notes |
|---|---|---|
HTTPS_PROXY, NO_PROXY, AIOHTTP_TRUST_ENV | HTTP clients / LiteLLM | See Proxies |
SSL_CERT_FILE, AWS_CA_BUNDLE | LiteLLM / AWS SDK | See TLS inspection and private CAs |
AWS_REGION_NAME, AWS_REGION | LiteLLM | Bedrock region when aws_region_name is not in FILE_SERVICE_AI__EXTRA_PARAMS. Prefer EXTRA_PARAMS in Option A; in Option B this is the only way to set the region, so set it explicitly (pod identity does not guarantee it) |
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN | AWS SDK | Static credentials for Bedrock. Prefer pod identity |
AWS_BEARER_TOKEN_BEDROCK | LiteLLM | Fallback Bedrock API key when FILE_SERVICE_AI__API_KEY is unset. Prefer FILE_SERVICE_AI__API_KEY |
Frontend service (istari-frontend): VITE_EXPERIMENTAL_DEFAULT_ON, VITE_EXPERIMENTAL_LOCKED — see Frontend settings.
Network requirements
Egress from the registry service pods
| Destination | Port | Purpose |
|---|---|---|
| Your LLM provider hostname (see Provider setup) | 443 | Chat completions (long-lived streaming responses) |
mcp.<customer_istari_fqdn> | 443 | Tool calls to the MCP Service. This is your own FQDN, so in-cluster traffic must be able to reach it (hairpin through the load balancer, or split DNS) |
The registry service ships with LiteLLM's model catalog bundled and does not download it at startup, so no egress to GitHub or other LiteLLM hosts is needed.
The MCP Service in turn needs its own egress, which depends on whether the API Gateway is enabled:
- API Gateway enabled (
apiGateway.enabled: true): MCP callsapi.<customer_istari_fqdn>on port 443 at the/registryand/identityprefixes. - No API Gateway: MCP calls
registry.<customer_istari_fqdn>andzitadel.<customer_istari_fqdn>directly on port 443.
In both cases the MCP Service itself stays at mcp.<customer_istari_fqdn>; the gateway has no /mcp prefix. These routes already exist if MCP works for external MCP clients.
Firewalls, WAFs, and SNI allow-lists
- Allow the provider hostname by SNI/hostname, not by IP; provider IPs change.
- Chat replies stream over Server-Sent Events from
POST /api/v3/ai/chat/streamon the registry service (/registry/api/v3/ai/chat/streambehind the API Gateway). Anything between the browser and the registry (load balancer, WAF, reverse proxy) must not buffer responses and needs an idle timeout of several minutes. - The registry service disables HTTP redirects on LLM calls. A proxy that answers with a
30xredirect breaks the call rather than being followed.
Proxies
The registry service honors the standard HTTPS_PROXY and NO_PROXY variables for most providers. LiteLLM's default transport (used for Bedrock and a few other providers) additionally requires AIOHTTP_TRUST_ENV: "true". Put the cluster-internal hostnames the registry service talks to (PostgreSQL, SpiceDB, NATS, and mcp.<customer_istari_fqdn> if it resolves internally) in NO_PROXY.
Validate proxy settings in a staging environment first; proxy behavior differs by product and by provider adapter.
TLS inspection and private CAs
If outbound traffic passes through TLS inspection, or your provider gateway presents a certificate from a private CA, the registry service must trust that CA or every LLM call fails with CERTIFICATE_VERIFY_FAILED.
- Preferred: add the CA to the Helm chart's
trustedCertBundleas in Scenario 2: Self-Signed TLS Certs. - If AI calls still fail, point both
SSL_CERT_FILE(used by LiteLLM) andAWS_CA_BUNDLE(used by the AWS SDK) at one PEM bundle that contains the public roots and your inspection CA. A bundle that contains only your private CA breaks object-store and other public TLS calls from the same pod.
Validate the deployment
Run these from a workstation with kubectl access to the platform namespace. Replace <...> placeholders.
1. The registry service sees the configuration
kubectl exec deployment/istari-fileservice -- sh -c 'env | grep -E "^FILE_SERVICE_AI" | sed "s/=.*/=<set>/"'
FILE_SERVICE_AI__MODEL (or at least one FILE_SERVICE_AI_NN_HOST) and FILE_SERVICE_AI__MCP_SERVER_URL should be listed.
2. Outbound connectivity and TLS from the pod
Probe the provider you configured. Any HTTP status means the pod reached it over TLS; 000 means blocked egress or a TLS failure.
| Provider | Test URL | Expected |
|---|---|---|
| Anthropic | https://api.anthropic.com/v1/models | 401 |
| OpenAI | https://api.openai.com/v1/models | 401 |
| Azure OpenAI | https://<resource>.openai.azure.com/openai/models?api-version=2024-02-01 | 401 |
| AWS Bedrock | https://bedrock-runtime.<region>.amazonaws.com/ | 403 or 404 |
| OpenAI-compatible gateway | https://<gateway-host>/v1/models | 401, or 200 if it allows anonymous listing |
kubectl exec deployment/istari-fileservice -- curl -sS -o /dev/null -w '%{http_code}\n' "<test-url>"
3. Check for TLS inspection
kubectl exec deployment/istari-fileservice -- curl -sv -o /dev/null "https://<provider-host>" 2>&1 | grep -E 'issuer|SSL certificate'
The issuer should be a public CA. A corporate CA name means traffic is inspected; see TLS inspection and private CAs. SSL certificate problem means the inspecting CA is not trusted yet.
4. The MCP Service is reachable from the pod
kubectl exec deployment/istari-fileservice -- curl -sS -o /dev/null -w '%{http_code}\n' https://mcp.<customer_istari_fqdn>/mcp/
Expected: 401 (reachable; the request carried no token). 404 means the request reached your ingress but there is no route to the MCP Service at /mcp/. A 5xx means the MCP Service is reachable but failing; check the istari-mcp pod logs. 000 means the registry service pod cannot resolve or connect to the MCP FQDN.
5. The AI Chat endpoint answers
In the web app: turn on Istari AI Chat under Application → Experimental Features, click AI chat in the header, and send Reply with the single word OK. A reply confirms the provider path. Then ask What is the name of the system I am looking at? from a system page; a real answer confirms the MCP path.
The command-line test below calls the blocking POST /api/v3/ai/chat endpoint, which reads only the environment variables (Option A). On an Option B-only deployment it returns 404 even though chat works in the web app; use the browser test above instead.
From the command line, you need a bearer token for a platform user. With the Identity Service enabled the platform issues Keys rather than pasteable tokens, so let the Istari Python Client mint the token from a Key you generate under avatar → Developer → Generate Key (see Settings):
export ISTARI_DIGITAL_API_URL="<API URL shown under Developer → Endpoints>"
export ISTARI_CLIENT_IDENTITY_SERVICE_SECRET_FILE="/path/to/downloaded-key.json"
export ISTARI_DIGITAL_IDENTITY_SERVICE_ENABLED=true
AUTH=$(python -c 'from istari_digital_client import Configuration; print(next(v["value"] for v in Configuration().auth_settings().values()))')
curl -sS -X POST "<registry-base-url>/api/v3/ai/chat" \
-H "Authorization: $AUTH" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Reply with the single word OK."}]}'
<registry-base-url> is https://registry.<customer_istari_fqdn> for a direct deployment, or https://api.<customer_istari_fqdn>/registry behind the API Gateway. Where a Personal Access Token still exists, -H "Authorization: Bearer <pat>" works too.
200with amessageobject — configured and working.404withThe AI integration is not configured…—FILE_SERVICE_AI__MODELis unset (see step 1); expected on an Option B-only deployment.5xx— the provider call failed; read the registry logs (step 6).
6. Registry service logs
kubectl logs deployment/istari-fileservice --since=15m | grep -i -E 'litellm|ai_service|mcp|CERTIFICATE'
Most provider-side failures reach the user as a generic "Istari AI hit an error" message; the specific cause is only in these logs.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Toggle is on but the AI chat button is missing, or a toast says Istari AI is not configured on this server | No FILE_SERVICE_AI__MODEL and no enabled catalog model; or env vars set but the pod not restarted | Set the model (Option A) or enable a model (Option B); kubectl rollout restart deployment/istari-fileservice after env changes |
| Assistant chats but cannot see systems or files; replies contain text that looks like tool calls | FILE_SERVICE_AI__MCP_SERVER_URL unset, wrong path (must end in /mcp/), MCP not deployed, or the registry cannot reach mcp.<fqdn> | Set the URL, run validation step 4, check istari-mcp pod logs; fix hairpin/split DNS so pods reach your own FQDN |
Registry logs show CERTIFICATE_VERIFY_FAILED / unable to get local issuer certificate | TLS inspection or a private CA on the provider path | Run validation step 3; add the CA via trustedCertBundle, or set SSL_CERT_FILE and AWS_CA_BUNDLE to a full bundle |
Connection timeouts or 000 from validation step 2 | Egress blocked (security group, NAT, firewall, WAF SNI list) or proxy not configured | Allow the provider FQDN on 443; set HTTPS_PROXY/NO_PROXY/AIOHTTP_TRUST_ENV if a proxy is mandatory |
Provider returns 401/403 (in registry logs) | Wrong key for the provider (sk-ant- vs sk-), key for another resource/region, revoked key | Re-check the key and provider match; for Azure confirm the key belongs to the resource in BASE_URL |
Bedrock AccessDeniedException | Pod role lacks bedrock:InvokeModel*, model access not enabled in that region, marketplace subscription missing | Apply Bedrock IAM permissions; enable model access in the Bedrock console for that region; subscribe to the model in AWS Marketplace |
Bedrock ValidationException / model not found | Wrong model id or region; cross-region models need the us./eu. inference-profile id | Use the inference profile id (e.g. bedrock/us.anthropic.claude-opus-4-8); set aws_region_name (Option A) or the pod's AWS_REGION (Option B) to a region that serves it |
Azure 404 / DeploymentNotFound | FILE_SERVICE_AI__MODEL uses the model name instead of the deployment name, or API_VERSION missing | Use azure/<deployment-name>; set FILE_SERVICE_AI__API_VERSION |
Provider 400 mentioning temperature or an unsupported parameter for a brand-new model | The platform's bundled LiteLLM predates the model | Pick a model the current release supports, or upgrade the platform |
| AI Models page: Host picker is empty | No FILE_SERVICE_AI_NN_HOST/_AUTHN pairs on the registry service, or _AUTHN missing (not empty) | Provision slots per Option B and restart the registry service |
| Endpoint badge Host changed / Host not provisioned / Invalid host config | Slot hostname changed, slot removed, or _HOST value fails validation | Restore or fix the slot value; an administrator re-saves the endpoint to confirm |
| Replies stop mid-sentence; "The connection dropped before the reply finished" | A proxy or load balancer buffers or times out the SSE stream | Disable response buffering and raise idle timeouts on the path to /api/v3/ai/chat/stream |
Provider 429 | Provider rate limit or quota | Raise the quota with the provider or choose a different model |
Bedrock IAM permissions
This section is being finalized with Istari Digital's infrastructure team. Validate it against your account before relying on it.
To call Bedrock with the pod's IAM identity (no API key):
1. Enable model access for the models you want in the Bedrock console for the region you set in aws_region_name. Third-party models (Anthropic, Meta, and others) also require an active AWS Marketplace subscription in the account.
2. Attach a policy to the IAM role the registry service pods assume (EKS Pod Identity or IRSA). The template below covers one configured model id, <model-id>, for example anthropic.claude-opus-4-8 or amazon.nova-pro-v1:0. Replace <partition> (aws, or aws-us-gov in GovCloud), <region> (the region you call), <account-id>, <geo>, and the destination regions, then remove the resources that do not apply to how you invoke the model.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "BedrockInvokeModel",
"Effect": "Allow",
"Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
"Resource": [
"arn:<partition>:bedrock:<region>::foundation-model/<model-id>",
"arn:<partition>:bedrock:<region>:<account-id>:inference-profile/<geo>.<model-id>",
"arn:<partition>:bedrock:<destination-region-1>::foundation-model/<model-id>",
"arn:<partition>:bedrock:<destination-region-2>::foundation-model/<model-id>"
]
},
{
"Sid": "BedrockMarketplaceViewSubscriptions",
"Effect": "Allow",
"Action": ["aws-marketplace:ViewSubscriptions"],
"Resource": "*"
}
]
}
-
Direct model id (
bedrock/<model-id>, no geographic prefix): only the first resource is needed, the foundation model in the region you call. -
Cross-region inference profile (
bedrock/<geo>.<model-id>, where<geo>isus,eu,apac, and so on): needs the profile ARN in the region you call and a foundation-model ARN for every region the profile routes to. List them with:aws bedrock get-inference-profile --region <region> --inference-profile-identifier <geo>.<model-id> --query 'models[].modelArn' -
The template names one model. Repeat the resources for each model id you configure, or use a prefix wildcard where your policy allows, for example
arn:<partition>:bedrock:*::foundation-model/anthropic.claude-*for every region and Claude model. -
aws-marketplace:ViewSubscriptionslets the invoking principal verify the account's subscription for marketplace-gated models (Anthropic and other third-party models). Subscribing itself is a one-time account action, not something the pod role needs.
3. Optional: keep traffic inside the VPC with interface endpoints for bedrock-runtime (and bedrock) with private DNS enabled. No registry service change is needed; the same hostname resolves to the endpoint.
If your platform cannot use pod identity, create an IAM user with the same policy and either generate a Bedrock API key for it (set as FILE_SERVICE_AI__API_KEY, or as a slot's _AUTHN) or use the OpenAI-compatible endpoint described under AWS Bedrock.
Related
- Istari AI Chat — the end-user feature.
- Istari Platform Installation — secrets, MCP Service, trusted certificate bundles, custom environment variables.
- MCP quick start — the MCP Service and the LLM integration disclaimer.
- AI Models (Admin Guide) and AI Steering (Admin Guide) — the organization administrator screens.
- Granting admin access — organization administrator rights for the AI Models page.