Skip to main content
cloro offers two ways to call the API: synchronous requests, where you wait for the result on the same connection, and asynchronous requests, where you submit a task and retrieve the result later.

Synchronous requests

Synchronous endpoints return the full result in the same HTTP response, so the request has to fit inside your client’s timeout. See Synchronous requests for the request and response structure, common parameters, optional response formats, and code examples.

Asynchronous requests

Asynchronous endpoints accept a task, return a taskId immediately, and process the work in the background. You retrieve the result via webhook (recommended) or by polling. See Asynchronous requests for the two-step submit-and-fetch flow, task states, queue limits, and request prioritization.

When to use which

How a request runs

cloro does not call a vendor API for the AI surfaces. There is no public API behind ChatGPT, Perplexity, Copilot, Gemini, Grok, AI Overview or AI Mode that returns what a real user sees, so each request drives a real browser session through a proxy in the country you asked for: load the surface, submit the prompt, wait for the answer to finish streaming, then parse the rendered page into the response fields. A single prompt often needs several page loads, for example when the surface defers its sources or its product cards until the answer settles. That is where the seconds go, and it is also what makes the output match what a person in that country would see.

Common questions

How long should one request take?

Typical end-to-end times, measured over the last seven days of production traffic: A Google SERP is one page load, so it lands in seconds. An AI surface has to generate its answer before there is anything to read, and that generation time is the bulk of the wait. The spread moves with the provider’s own load, so treat these as the shape rather than a guarantee. Elmo publishes an independent comparison of the providers running these engines, re-measured every six hours, with cloro alongside other providers on the same prompts.

Why are my requests slower than expected?

Beyond the generation time above, latency comes from:
  • Queue depth (async requests only): set by your plan’s concurrency limit and your current queue size
  • Geographic routing: requests route through region-specific infrastructure, which may differ in latency depending on the country parameter you specify
  • Provider load: peak hours on the provider’s side lengthen generation
If a batch takes far longer than one request times the number of prompts, the cause is almost always concurrency rather than any single request: on one slot the prompts run one after another. Upgrading your plan raises concurrency, which drains the async queue faster. To measure a single request, see Can I read the request latency from the API response?.

Can I get faster processing for high-volume workloads?

Yes — the larger the plan, the greater the concurrency assigned to it. For async requests, higher concurrency means more tasks processed at once and shorter queue waits. For sync requests, it means you can send more simultaneous requests without hitting limits.

How do I submit many requests in one call?

Use the batch endpoint. It takes up to 500 tasks in a single HTTP request, so you avoid the overhead of 500 round trips, and it validates each task independently — one bad task doesn’t block the rest. You can hold up to 100,000 tasks in the queue at a time, and cloro handles queuing, so you don’t manage concurrency limits yourself. Retrieve results with webhooks or polling. For which pattern to use once you’re submitting at volume — async with webhooks, or sync with concurrent workers — see the concurrency guide, which has code examples for both.