Production checklist
Lessons from running the package in several production applications. Each item is something that went wrong at least once.
Queues
Section titled “Queues”- A dedicated queue connection for
aiwithretry_afterabove the job timeout. On the defaultredisconnection (retry_after90) a provider call running longer than 90 seconds is handed out again, and the user gets two replies. TheQueuerow ofphp artisan aboutwarns about it. See Queues and Horizon. ai-postin a shared pool of short jobs — it rarely has work, a separate supervisor just holds an idle process.- Watch the job timeout budget. A queued run may make several provider calls in a row (a fallback chain, a forced second pass), each with its own HTTP timeout, plus the time of the tools. Keep the sum below
jobTimeout(). - Fan-out vs. one job. The prompt is built when a task is queued. A batch that dispatches many
AI::queue()calls at once builds all prompts from the same snapshot — if each prompt should see the results of the previous ones (e.g. duplicate detection against already processed items), runAI::send()sequentially inside one queued job instead, with a batch size that fits the job timeout,withoutOverlapping()and a lock.
Failover
Section titled “Failover”-
Configure a fallback chain from the environment, so a second provider is one
.envline away:$chain = array_values(array_unique(array_filter([env('AI_DEFAULT', 'openai'), env('AI_FALLBACK')])));'routing' => ['chat_reply' => $chain,'summarize' => $chain,],Routing keys are task names — a task missing from
routinggets only the default driver and no fallback. -
Lower the HTTP timeout for interactive tasks (
options['timeout'] => 30): a hung provider otherwise holds the request for the default 60 seconds before the next driver is tried. -
Don’t list drivers without API keys in
viaDrivers()orrouting: each one leaves askippedrow (driver_not_configured) inai_runson every sync call. -
Implement
onFailed()on user-facing tasks. Without it a provider outage ends silently — no reply and nobody notified. See TheonFailed()hook. -
Check the driver state on the dashboard: an outage covered by the fallback is visible only there. With several servers use a shared cache store.
Security
Section titled “Security”-
Protect the dashboard:
'middleware' => ['web', 'auth', 'can:admin']and a non-default path such asadmin/ai-tasks. The default['web']exposes prompts, responses and the Retry / Dead actions. -
Store per-tenant API keys encrypted (
'api_key' => 'encrypted'cast) and pass them viaproviderOverrideonly when set:public function providerOverride(): ?array{if (empty($this->api_key)) {return null;}return array_filter(['driver' => $this->driver, 'key' => $this->api_key, 'model' => $this->model]);} -
Treat model output as untrusted. HTML generated by a model and rendered with
{!! !!}or into a PDF needs sanitizing; check$response->okbefore using the content. -
Keep
store_requestoff unless you needai:retry/ dashboard Retry — prompts often contain personal data.
Switches and limits
Section titled “Switches and limits”- A global kill switch in your own config (
AI_ENABLED), checked before dispatch and inshouldRun(). It also keeps tests and seeders from calling providers. - Per-feature toggles checked in
shouldRun(), which runs on the worker — then a feature switched off while jobs wait in the queue takes effect immediately.toPayload()andtools()run at dispatch, so a toggle read there applies only to newly queued tasks. - Rate-limit automatic replies per user/chat (
RateLimiter), and hand over to a human when the limit is hit — a loop with another bot otherwise burns tokens indefinitely.
Tenants and billing
Section titled “Tenants and billing”- Override
tenantId()on every task. A queued job has neither theX-Tenant-Idheader nor an authenticated user, so the default resolver puts every queued run underdefault. Even sync calls can resolve wrongly: a manager calling the AI on behalf of a customer’s organization is not the tenant. See Budgets & tenants. - Set
subjectType()/subjectId()— filteringai_runsby the chat or record a run belongs to is the first thing needed when investigating a complaint. - Charging your users per token: listen to
AiRunFinishedand chargetokens_in + cache_read_tokens + cache_write_tokens + tokens_out.tokens_inexcludes cached input, so charging only it would make the user’s price depend on whether the provider cache hit. Use the run id as the idempotency key of the charge, and guard against non-tenant ids —ai:requestwithout--tenantrecordsdefault.
- List every model under
prices, including aliases (deepseek-flashanddeepseek-v4-flash). A model that is not listed falls back to the driver’sprice, visible ascost_rates.source = "driver"— query for it after each model switch. - Re-check the published config after upgrading.
composer updatenever touchesconfig/ai-tasks.php; new keys and removed options are listed in theConfigrow ofphp artisan about, a missing migration inSchema. Changed default models and prices are not detected — compare those by hand. - Known gaps: DeepSeek peak-hour pricing and TTS costs are approximations, see Cost tracking.
Upgrades
Section titled “Upgrades”- Run an evaluation before changing the model, prompt, schema, history format or a package version — see Evaluating prompts. Provider behaviour shifts in ways no unit test catches.
- Under Octane a sync endpoint is limited by the request timeout: cap
max_tokensfor long generations, or queue them.