Skip to content

Production checklist

Lessons from running the package in several production applications. Each item is something that went wrong at least once.

  • A dedicated queue connection for ai with retry_after above the job timeout. On the default redis connection (retry_after 90) a provider call running longer than 90 seconds is handed out again, and the user gets two replies. The Queue row of php artisan about warns about it. See Queues and Horizon.
  • ai-post in a shared pool of short jobs — it rarely has work, a separate supervisor just holds an idle process.
  • Watch the job timeout budget. A queued run may make several provider calls in a row (a fallback chain, a forced second pass), each with its own HTTP timeout, plus the time of the tools. Keep the sum below jobTimeout().
  • Fan-out vs. one job. The prompt is built when a task is queued. A batch that dispatches many AI::queue() calls at once builds all prompts from the same snapshot — if each prompt should see the results of the previous ones (e.g. duplicate detection against already processed items), run AI::send() sequentially inside one queued job instead, with a batch size that fits the job timeout, withoutOverlapping() and a lock.
  • Configure a fallback chain from the environment, so a second provider is one .env line away:

    $chain = array_values(array_unique(array_filter([env('AI_DEFAULT', 'openai'), env('AI_FALLBACK')])));
    'routing' => [
    'chat_reply' => $chain,
    'summarize' => $chain,
    ],

    Routing keys are task names — a task missing from routing gets only the default driver and no fallback.

  • Lower the HTTP timeout for interactive tasks (options['timeout'] => 30): a hung provider otherwise holds the request for the default 60 seconds before the next driver is tried.

  • Don’t list drivers without API keys in viaDrivers() or routing: each one leaves a skipped row (driver_not_configured) in ai_runs on every sync call.

  • Implement onFailed() on user-facing tasks. Without it a provider outage ends silently — no reply and nobody notified. See The onFailed() hook.

  • Check the driver state on the dashboard: an outage covered by the fallback is visible only there. With several servers use a shared cache store.

  • Protect the dashboard: 'middleware' => ['web', 'auth', 'can:admin'] and a non-default path such as admin/ai-tasks. The default ['web'] exposes prompts, responses and the Retry / Dead actions.

  • Store per-tenant API keys encrypted ('api_key' => 'encrypted' cast) and pass them via providerOverride only when set:

    public function providerOverride(): ?array
    {
    if (empty($this->api_key)) {
    return null;
    }
    return array_filter(['driver' => $this->driver, 'key' => $this->api_key, 'model' => $this->model]);
    }
  • Treat model output as untrusted. HTML generated by a model and rendered with {!! !!} or into a PDF needs sanitizing; check $response->ok before using the content.

  • Keep store_request off unless you need ai:retry / dashboard Retry — prompts often contain personal data.

  • A global kill switch in your own config (AI_ENABLED), checked before dispatch and in shouldRun(). It also keeps tests and seeders from calling providers.
  • Per-feature toggles checked in shouldRun(), which runs on the worker — then a feature switched off while jobs wait in the queue takes effect immediately. toPayload() and tools() run at dispatch, so a toggle read there applies only to newly queued tasks.
  • Rate-limit automatic replies per user/chat (RateLimiter), and hand over to a human when the limit is hit — a loop with another bot otherwise burns tokens indefinitely.
  • Override tenantId() on every task. A queued job has neither the X-Tenant-Id header nor an authenticated user, so the default resolver puts every queued run under default. Even sync calls can resolve wrongly: a manager calling the AI on behalf of a customer’s organization is not the tenant. See Budgets & tenants.
  • Set subjectType() / subjectId() — filtering ai_runs by the chat or record a run belongs to is the first thing needed when investigating a complaint.
  • Charging your users per token: listen to AiRunFinished and charge tokens_in + cache_read_tokens + cache_write_tokens + tokens_out. tokens_in excludes cached input, so charging only it would make the user’s price depend on whether the provider cache hit. Use the run id as the idempotency key of the charge, and guard against non-tenant ids — ai:request without --tenant records default.
  • List every model under prices, including aliases (deepseek-flash and deepseek-v4-flash). A model that is not listed falls back to the driver’s price, visible as cost_rates.source = "driver" — query for it after each model switch.
  • Re-check the published config after upgrading. composer update never touches config/ai-tasks.php; new keys and removed options are listed in the Config row of php artisan about, a missing migration in Schema. Changed default models and prices are not detected — compare those by hand.
  • Known gaps: DeepSeek peak-hour pricing and TTS costs are approximations, see Cost tracking.
  • Run an evaluation before changing the model, prompt, schema, history format or a package version — see Evaluating prompts. Provider behaviour shifts in ways no unit test catches.
  • Under Octane a sync endpoint is limited by the request timeout: cap max_tokens for long generations, or queue them.