Synchronous requests
Synchronous endpoints return the full result in the same HTTP response, so the request has to fit inside your client’s timeout. See Synchronous requests for the request and response structure, common parameters, optional response formats, and code examples.Asynchronous requests
Asynchronous endpoints accept a task, return ataskId immediately, and process the work in the background. You retrieve the result via webhook (recommended) or by polling. See Asynchronous requests for the two-step submit-and-fetch flow, task states, queue limits, and request prioritization.
When to use which
How a request runs
cloro does not call a vendor API for the AI surfaces. There is no public API behind ChatGPT, Perplexity, Copilot, Gemini, Grok, AI Overview or AI Mode that returns what a real user sees, so each request drives a real browser session through a proxy in the country you asked for: load the surface, submit the prompt, wait for the answer to finish streaming, then parse the rendered page into the response fields. A single prompt often needs several page loads, for example when the surface defers its sources or its product cards until the answer settles. That is where the seconds go, and it is also what makes the output match what a person in that country would see.Common questions
How long should one request take?
Typical end-to-end times, measured over the last seven days of production traffic:
A Google SERP is one page load, so it lands in seconds. An AI surface has to
generate its answer before there is anything to read, and that generation time
is the bulk of the wait. The spread moves with the provider’s own load, so
treat these as the shape rather than a guarantee.
Elmo publishes an independent comparison of the
providers running these engines, re-measured every six hours, with cloro
alongside other providers on the same prompts.
Why are my requests slower than expected?
Beyond the generation time above, latency comes from:- Queue depth (async requests only): set by your plan’s concurrency limit and your current queue size
- Geographic routing: requests route through region-specific infrastructure, which may differ in latency depending on the
countryparameter you specify - Provider load: peak hours on the provider’s side lengthen generation