Docs

Install the client, or call the endpoint.

Almost everyone wants the first one: our source-available solver drives the page and uses this endpoint for you. Underneath, CaptchaKraken speaks the OpenAI chat-completions API — so if your tooling can talk to OpenAI, it can talk to this.

Setup

Base URL and key.

# same client, same request shape — only the endpoint moves
export CAPTCHA_KRAKEN_API_KEY=ck_live_…
export VLLM_BASE_URL=https://api.captchakraken.com/v1

Keys are minted from the dashboard, or by the MCP server below. They look like ck_live_ followed by a public prefix and a secret half, and the secret is shown exactly once at issue. We store only a SHA-256 of it, so a lost key is replaced rather than recovered.

Authenticate with a standard bearer header: Authorization: Bearer ck_live_…. The endpoint is api.captchakraken.com.

Quickstart

Point the client at a page.

CaptchaKraken is our public solver. It finds the widget, screenshots it, sends the right prompt, clicks the answer and checks whether the vendor accepted it. Give it a key and it uses this endpoint — you write none of the request shape below.

npm install captchakraken    # TypeScript browser driver
pip install captchakraken    # Python engine + the `captchakraken` CLI

export CAPTCHA_KRAKEN_API_KEY=ck_live_your_key_here

TypeScript — any Playwright-compatible page

import { chromium } from 'playwright';
import { CaptchaKrakenSolver } from 'captchakraken';

const page = await (await (await chromium.launch()).newContext()).newPage();
await page.goto('https://www.google.com/recaptcha/api2/demo');

await new CaptchaKrakenSolver().solve(page);   // detect → solve → click → verify

Python

from playwright.sync_api import sync_playwright
from captchakraken import PageSolver

with sync_playwright() as p:
    page = p.chromium.launch().new_page()
    page.goto("https://www.google.com/recaptcha/api2/demo")
    print(PageSolver().solve(page).is_solved)

No browser? Hand the CLI a screenshot and it prints the click plan as JSON: captchakraken path/to/captcha.png. Every browser framework it drives — Playwright, Patchright, Camoufox, Puppeteer — is in the repository's usage guide.

The rest of this page is the wire protocol underneath it. You need it to write your own client or to debug one — not to solve a captcha.

Important

This is a captcha model, not a general vision model.

The endpoint accepts three prompts — grid selection, click/drag, and the animated-challenge prompt — and rejects anything else with a 400 unrecognized_prompt. It is not a general-purpose image endpoint and will not behave as one.

Send the prompts as written. They are pinned to the model's training distribution, and paraphrasing them measurably costs accuracy.

Request

Solving a grid challenge.

Send the captcha screenshot as a data URL and the prompt as text, in that order, inside one user turn.

curl https://api.captchakraken.com/v1/chat/completions \
  -H "Authorization: Bearer $CAPTCHA_KRAKEN_API_KEY" \
  -H "Content-Type: application/json" \
  -H "x-ck-session: $(uuidgen)" \
  -d '{
    "messages": [
      {
        "role": "system",
        "content": "You are an expert captcha solver. Respond ONLY with the JSON action."
      },
      {
        "role": "user",
        "content": [
          { "type": "image_url",
            "image_url": { "url": "data:image/png;base64,<SCREENSHOT>" } },
          { "type": "text", "text": "<PROMPT — see below>" }
        ]
      }
    ],
    "temperature": 0,
    "max_tokens": 128,
    "chat_template_kwargs": { "enable_thinking": false }
  }'

There is no model to choose

The example above sends no model field, and that is not an omission — you do not need one. There is a single hosted model and we point it at whichever weights are current, so when we retrain, your integration picks up the better model without you changing a line or knowing it happened.

OpenAI-compatible SDKs are the exception, and only because model is a required argument in those libraries rather than something we ask for. Send "captcha" and forget about it — any value you send resolves to the same place, and GET /v1/models will only ever list that one name.

The three prompts, verbatim

Only if you are building your own client — the CaptchaKraken package above already sends the right one per challenge. They are pinned to the model's training distribution: send them exactly as written, because paraphrasing measurably costs accuracy, and anything unrecognised is a 400 unrecognized_prompt.

Grid selection prompt

Substitute the real grid size and the matching hint line — Separate images. for a 3×3 of distinct photos, Single large image split into tiles. for a 4×4 subdivided one.

Solve the captcha grid by choosing the cell numbers that match the description from the captcha image prompt.

Grid: 3x3 (9 cells)
Hint: Separate images. Select only clear matches.

A cell that is already selected — small checkmark badge, border or highlight — still counts. Include it if it matches.

A cell being REPLACED does not: a large checkmark over the middle of the picture, a picture fading to white, or a new picture fading in. That cell is on its way to showing something else, so leave it out however well it matches.

If no cells match the description, return an empty list for target_ids: [].

Return JSON Array: [list of cell numbers (1-9)]
Click / drag prompt

For any challenge that is not a grid — click the matching icons in order, drag a piece into its slot, trace a path. No substitutions. Coordinates come back on a normalised 0–1000 scale, so they survive whatever size you screenshotted at.

Your task is to solve the captcha. Read the instruction at the top of the image carefully.

Look at the puzzle and decide what action solves it. All coordinates you return must be on a normalized 0–1000 image scale (top-left = (0, 0), bottom-right = (1000, 1000)).

Name WHAT each object is with a short 1–2 word label, then give its position. Choose ONE response:

FOR CLICK PUZZLES:
  Label each thing you click and give its point — subjects[i] names points[i]:
  → "action": "click", "subjects": ["<label>", ...], "points": [[x1, y1], [x2, y2], ...]

FOR DRAG PUZZLES:
  Drag ONE item at a time. Label the source (the piece you pick up) and the destination (where it belongs), each with a short 1–2 word label, and give both points:
  → "action": "drag", "drags": [{ "source": "<label>", "from": [x, y], "destination": "<label>", "to": [x, y] }, ...]

FOR PUZZLE PIECE SLIDER PUZZLES:
  A single jigsaw piece has to end up in the piece-shaped slot cut into the picture. Do not pick up the piece or the slider handle — leave the source EMPTY and give only the destination, the CENTER OF THE SLOT the piece belongs in:
  → "action": "drag", "drags": [{ "source": "", "from": [], "destination": "<label>", "to": [x, y] }]

Respond ONLY with JSON:
{
  "action": "click", "subjects": [ ... ], "points": [ ... ]
  // OR "action": "drag", "drags": [ ... ]
}
Animated-challenge prompt

For challenges that move. Send several keyframes cut from a short recording instead of one screenshot, as separate image parts in the same message. Substitute the number of frames you sent: below is the four-frame form, and the listing is simply frame 1, frame 2, … up to that count. Animated challenges bill at the video rate.

Your task is to solve the captcha. This challenge is animated, so instead of one picture you are given 4 still keyframes cut from a short recording of it, in order: frame 1, frame 2, frame 3, frame 4.

Every keyframe shows the SAME puzzle at a different moment. Read the instruction at the top of the keyframes carefully. What you need to act on may be visible in only some of the frames — sprites fade in and out, boards cycle their contents — so pick the ONE frame in which your target is clearest and report its number as "frame".

The frame number is there so the solver knows WHEN to press the mouse: it waits for the widget to look like that frame before clicking. If your answer does not depend on the frame — the target is in the same place in every one of them — then there is nothing to wait for, and you may leave "frame" out entirely.

Read your coordinates off THAT frame. All coordinates must be on a normalized 0–1000 image scale (top-left = (0, 0), bottom-right = (1000, 1000)).

Name WHAT each object is with a short 1–2 word label, then give its position. Choose ONE response:

FOR CLICK PUZZLES:
  Label each thing you click and give its point — subjects[i] names points[i]:
  → "frame": <1-4>, "action": "click", "subjects": ["<label>", ...], "points": [[x1, y1], [x2, y2], ...]

FOR DRAG PUZZLES:
  Drag ONE item at a time. Label the source (the piece you pick up) and the destination (where it belongs), each with a short 1–2 word label, and give both points:
  → "frame": <1-4>, "action": "drag", "drags": [{ "source": "<label>", "from": [x, y], "destination": "<label>", "to": [x, y] }, ...]

Respond ONLY with JSON:
{
  "frame": <1-4>,   // omit if your answer holds in every frame
  "action": "click", "subjects": [ ... ], "points": [ ... ]
  // OR "action": "drag", "drags": [ ... ]
}

From an OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CAPTCHA_KRAKEN_API_KEY"],
    base_url="https://api.captchakraken.com/v1",
)

response = client.chat.completions.create(
    model="captcha",
    messages=[...],                      # shape as above
    temperature=0,
    max_tokens=128,
    extra_headers={"x-ck-session": session_uuid},
)

The response is a normal chat completion

The answer arrives as JSON in the message content. Grid challenges return the cell numbers to click; click/drag challenges return labelled coordinates on a normalised 0–1000 scale.

{
  "choices": [
    { "message": { "role": "assistant", "content": "[6, 8]" } }
  ],
  "usage": { "prompt_tokens": 1102, "completion_tokens": 7 }
}

Headers

Two optional headers that affect your bill.

HeaderWhy it matters
x-ck-sessionA UUID identifying one captcha attempt. Responses sharing a session are grouped, and that grouping is what both per-attempt limits are counted over — the 5-response billing ceiling and the 10-response abandon threshold. Without it, every response bills independently and neither engages — send a fresh UUID per captcha, and reuse it across every response for it.
x-ck-clientIdentifies the integration, e.g. camoufox/0.4.11. Used for attribution and support; it has no effect on price.

There is no header that sets the puzzle class. Pricing is derived from the prompt in the body, so a header cannot move a request into a cheaper bracket.

Errors

Branch on the code, never on the message.

Errors use OpenAI's envelope, so your SDK will parse them. The prose gets reworded; error.code does not.

StatusCodeMeaningWhat to do
401missing_api_keyNo Authorization header was sent.Send Authorization: Bearer ck_live_…
401invalid_api_keyThe key is unknown or has been revoked.Check the key, or mint a new one.
402insufficient_creditsThe account balance cannot cover the round.Top up. The same key resumes working immediately.
403account_suspendedThe account is suspended.Contact support.
429rate_limitedToo many requests from this key.Back off; honour the Retry-After header.
400unrecognized_promptThe prompt is not one this service solves.Use one of the three supported prompts, unmodified.
400invalid_requestThe body is malformed.Check the JSON against the shape above.
413request_too_largeThe screenshot exceeds the body limit.Downscale or crop before encoding.
409solve_abandonedOne attempt was served 10 times without settling.Usually IP reputation, not the answers. Start a new x-ck-session.
502upstream_unavailableThe solver fleet is unreachable. Not your fault.Retry shortly with backoff. Nothing is billed.

A revoked key and an unknown key both return invalid_api_key. Telling the difference apart would confirm to an attacker which of their guesses was once real.

MCP

Managing the account from an agent.

The same account, driven from an MCP client. It signs you in through GitHub, mints and revokes keys, and reads back what you have spent.

# Claude Code, Claude Desktop, or any MCP client
claude mcp add captchakraken -- npx -y captchakraken-mcp

Sign in

sign_in prints a short code and a link. You approve it in the browser with GitHub; the account is created on the spot if it does not exist, free credits and all.

Mint keys

create_api_key returns a live ck_live_ key, once. list_api_keys and revoke_api_key do the rest.

Watch the money

get_balance and get_usage report the balance, the burn rate and the per-day spend. get_topup_link opens Stripe.

The token the MCP server holds manages the account — it cannot solve captchas, and an API key cannot manage the account. Two credentials, two blast radii. Revoke either from the dashboard.

Billing

$0.30 per 1,000 image responses.

You are billed per inference response — no subscription, no monthly minimum, no per-seat fee, and no charge at all for challenges that never reach the model.

What you sendCredits eachPer 1,000 responses
Image response
Grid, click and drag puzzles
3$0.30
Video response
Video challenges, which only the hosted model handles
10$1.00
reCAPTCHA checkbox, Cloudflare Turnstile
Detected locally; no model call is made.
0Free

Most image captchas are solved in 1–2 image inference requests, and most videos are solved in 1–2.

$1.00 = 10,000 credits. The per-response rate is the contractual figure and is what your account is actually charged; the response counts above are what we typically see, not a guarantee — a captcha that fights back takes more, and you are billed for what was actually served.

Credits

Buy what you need. It does not expire.

Any whole dollar amount from $5 up to $2,000— these three are shortcuts, not the only sizes. Amounts of $25 and over earn bonus credits, up to +50%; the table is on the top-up page.

$25262,500 credits

includes a +5% volume bonus

87,500 image responses

$50550,000 credits

includes a +10% volume bonus

183,333 image responses

$1001,180,000 credits

includes a +18% volume bonus

393,333 image responses

Those counts are division, not an estimate — 3 credits per image response into the pack, bonus credits included. Credits carry over indefinitely and are not tied to a billing period. Need more than $2,000 at once, or an invoice? Talk to us.

How billing works

Four rules, and no others.

One response, one charge

Every inference response costs the credits listed for its class. The class is derived from what you actually sent — an image or a video — not from a header, so a request cannot be relabelled into the cheaper bracket by anyone, including us.

Our misses are free

When the model under-selects and the challenge must be asked again for that reason, the extra response costs nothing and is labelled missed-tiles-retry on your usage so you can see it happened.

One captcha, one ceiling

At most 5 responses are billable per solve attempt — $1.50 per 1,000 even in the worst case. We keep answering free up to 10, then abandon the attempt. A pathological captcha cannot drain an account.

Failures we caused are not billed

If a response fails on our side — capacity, an error, anything that is not an answer — it is recorded at zero credits and appears on your usage with the reason.

Two things move your effective cost, and you should hear them here rather than infer them from an invoice. A reCAPTCHA grid re-draws after each click, so a session on a flagged IP draws harder challenges and lands above the typical response count. And if a vendor rejects a submission the model got right — a function of IP reputation and browser fingerprint more than of the answer — that response was still served and still billed. Tell us your setup and we will estimate it with you before you spend anything.

Common questions

Before you ask.

Is there a subscription?

No. Credits are prepaid and drawn down as you use them. There is no monthly fee, no seat count, and no minimum spend.

Do credits expire?

No. They stay on the account until they are used.

What if I run out mid-run?

Requests are refused with a clear insufficient_credits error rather than being served into a negative balance. Top up and the same key resumes working immediately.

Can I get a refund?

Unused credits, yes — see the refund policy. Credits already spent on solves we delivered are not refundable.

Do you offer volume pricing?

Yes, above the published packs. Get in touch with your expected monthly volume and challenge mix.