Skip to main content
See Making requests for an overview of both request modes. You submit a task, we return a taskId, and the request runs in the background — like a tracking number rather than waiting at the counter. Submit tasks one at a time or in batches of up to 500. Use async when:
  • Your application has short execution limits, typically in a serverless environment.
  • You want to submit many requests quickly without waiting on each one.
  • You need resilience that doesn’t depend on a single long-running connection.

Step 1: make an API request

Include a webhook.url and we notify you when the job is done; omit it and you poll for the result.
Identifying your requests with idempotencyKeyYou can optionally include an idempotencyKey in your request. This is a unique string you create that allows you to easily identify and reference your requests in your own system.The idempotencyKey must be a unique string across your entire account. If you submit a request with an idempotency key that has already been used, the API will return an error:
Generate unique keys using UUIDs, timestamps, or a combination of user ID + timestamp to ensure no duplicates.Failed task key retention: When a task returns FAILED, the idempotencyKey is not released — it stays bound for 24 hours, then auto-releases. To retry a failed task within that window, submit the same request with a new key.Serverless crash recovery: GET /v1/async/task/{taskId} finds a task by its taskId only, not by its idempotencyKey. Store the taskId as soon as you receive the submit response. If your serverless function (Vercel, Lambda) crashes before it saves the taskId, a resubmit with the same idempotencyKey returns 409 and does not create a second task, so you are not charged twice. Set a webhook.url when you submit: the webhook payload carries task.id and task.idempotencyKey, so you still receive the result.
We’ll acknowledge your request and provide a taskId.
Now, you can store this taskId and wait for the results.
Understanding task statesEvery async task goes through four possible states:
  • QUEUED: The request is received and waiting its turn to be processed. Tasks are processed by priority first, then in FIFO (first-in, first-out) order within the same priority level.
  • PROCESSING: The request is actively being processed by our system. The AI provider is generating your response.
  • COMPLETED: The request finished successfully. The final response is included in the same payload.
  • FAILED: The request failed to complete. This can happen because of rate limits, provider errors, or invalid input.
The initial response always shows status as QUEUED. You can track state transitions by polling the task status endpoint or receiving webhook updates.COMPLETED and FAILED tasks are stored in our system for 24 hours after completion. During this time, you can retrieve the full results using the task ID. After 24 hours, the task record and its associated response data are permanently deleted from our system.HTML URLs included in responses expire after 24 hours from generation, regardless of the task’s retention status.

Step 2: receive the results

Two ways to retrieve a result. Either way, a finished task carries task.latencyMs — the milliseconds from first pickup to the final outcome, excluding the initial queue wait. It is null while the task is QUEUED or PROCESSING, and stays null on a task that failed before processing ever started. On a retried task it covers every attempt, including the backoff between them, so it can be far longer than one provider call.
Don’t read task latency from the X-Latency-Ms header on a status call. That header times the status call itself — a few milliseconds — not the task. Task latency is only ever task.latencyMs.
If you provided a webhook.url in your request, we will send an HTTP POST to that URL containing the full result as soon as the task reaches a terminal state. See Webhooks for the payload shape, retry behavior, signature verification, and troubleshooting.

Option B: polling

Without a webhook URL, GET the task status endpoint with your taskId. Once the task completes, the response carries the full result.

Understanding limits

Task submission limits

At submission we check only two things:
  • Credit Limit: We verify you have enough credits for the task.
  • Queue Limit: Your organization can have a maximum of 100,000 tasks waiting in the queue. If you exceed this, you will receive a 429 Too Many Requests error. Please contact our team if you need this limit increased.

Task processing limits

Once your task is in the queue, our scheduler picks it up for processing. This is where your subscription’s concurrency limit is enforced. For example, if your plan allows 10 concurrent requests, our scheduler will process up to 10 of your tasks in parallel. Tasks are processed by priority first, then in the order they were received within the same priority level.

Request prioritization

Omitting it leaves the task at 1, so existing integrations are unaffected. Monitor how your queue is distributed across levels via the async status endpoint.
For practical concurrency patterns and examples, see our concurrency documentation.

Common questions

How do I cancel pending async tasks?

Individual queued tasks cannot be canceled by taskId — once a task enters the QUEUED state, it stays there until it is either processed (COMPLETED / FAILED) or removed by a queue-wide clear. To wipe every pending task in one call, use the DELETE /v1/async/queue endpoint. It removes all QUEUED tasks for your organization and returns the number cleared:
cURL
Behavior to be aware of:
  • Only QUEUED tasks are removed. PROCESSING tasks are already in-flight on a worker and keep running — they cannot be recalled.
  • COMPLETED and FAILED tasks are left in place and can still be retrieved via GET /v1/async/task/{taskId}.
  • Queued tasks have not been charged yet, so clearing the queue does not refund or debit credits.
  • The call is idempotent — re-running it on an empty queue returns cleared: 0.
Best practices to avoid unwanted tasks in the first place:
  • Test with small batches first
  • Use unique idempotencyKey values to prevent duplicate submissions
  • Implement safeguards in your submission logic
  • Monitor your queue depth via the async status endpoint before submitting large batches

What’s the maximum queue depth?

Queue depth is limited to 100,000 tasks per organization. If you exceed this limit, you’ll receive a 429 Too Many Requests error when trying to submit additional tasks. If you need a larger queue for your use case, please contact our team.

How long do async tasks stay in the queue?

Tasks stay in the queue until they are processed, resulting in either COMPLETED or FAILED status. There is no fixed time-to-result. How quickly a task completes depends on:
  • Queue depth: how many of your tasks are ahead of it. The scheduler runs higher priority tasks first, then FIFO within the same level.
  • Concurrency limit: your plan’s limit caps how many of your tasks process in parallel, so a deep queue drains at most that many tasks at a time. See Rate & concurrency limits.
  • Provider processing time: once a task is picked up, the time to the final outcome is reported in task.latencyMs. Queue wait is excluded from that number.
Check your current queue depth and concurrency usage with the async status endpoint.

How do I track the credits consumed by each task?

Both the polling response and the webhook payload include a credits object:
  • creditsToCharge — the estimated cost shown while the task is QUEUED or PROCESSING
  • creditsCharged — the actual amount billed once the task reaches COMPLETED or FAILED
Use creditsCharged from the terminal state to attribute cost per job. Failed tasks may still incur credits depending on how far processing got. Sync requests canceled by the client are also charged for work already done; async tasks cannot be canceled individually once queued, but you can wipe the whole pending queue in one call.

What happens to async tasks when my credits run out?

Credits are checked twice, and the balance is never reserved in advance:
  • At submission. POST /v1/async/task rejects with 403 INSUFFICIENT_CREDITS when your balance does not cover that task’s cost. POST /v1/async/task/batch instead returns 200 with a per-task INSUFFICIENT_CREDITS error for each task that doesn’t fit, so a partially affordable batch is partially accepted.
  • At scheduling. Tasks already sitting in QUEUED are re-checked when the scheduler picks them up. A task that no longer fits the balance moves to FAILED with creditsCharged: 0 and an INSUFFICIENT_CREDITS error — it is not held, retried, or resumed when you top up. Re-submit those tasks after topping up.
Both checks are skipped if your organization is on a subscription with overages enabled. Because credits are only deducted when a task completes, the balance from GET /v1/credits reflects completed charges only — it does not net out creditsToCharge for work still queued or processing. To decide whether a submission will fit, subtract your own outstanding creditsToCharge from remaining rather than trusting remaining alone.

A task came back COMPLETED but the result looks like an upstream error. What happened?

COMPLETED means cloro finished its work and returned what the upstream provider gave us. If the provider returned an error page, a captcha, or a truncated answer, that detail lives inside the response payload. Inspect the response body — FAILED is reserved for cases where cloro could not produce a result at all (network errors, internal exceptions, repeated upstream timeouts).

Does async cost more credits than sync?

No — async is 2 credits cheaper per request. The base cost and feature add-ons are identical whether you use synchronous or async delivery, but sync monitor requests (/v1/monitor/*) carry a +2 credit surcharge that async and batch requests (/v1/async/*) do not. The creditsCharged field in the terminal task state shows the actual amount billed for each task.

The async endpoint feels slow today. What should I check?

Call GET /v1/async/status to see your account’s queue depth and concurrency usage. If the queue is deep and concurrency is saturated, throughput is plan-bound — see concurrency for how to raise it.