# Date: 2026-08-23 # Author: Alok # File: docs/public/openapi.yaml # Purpose: Single machine-readable source of truth for the public DocXtract v3.1 API. # Generates the Postman collection, the portal API reference, and SDK types. # # Scope: api/v3.1 ONLY. v1 and v2 are retired (404). v3 is live but maintenance-only. # All shapes below were read from the v3.1 source, not from prose docs. openapi: 3.1.0 info: title: DocXtract API version: "3.1" summary: AI-powered document extraction — structured JSON from invoices, KYC, HR, and financial documents. description: | DocXtract extracts structured JSON from PDFs and images across 20+ document types. ## Base URL `https://api.docxtract.io/` There is **no `/api` prefix**. The `api/` directory in the repository is the docroot for the `api.` subdomain; it is not part of the public path. The endpoint is `/v3.1/documents`, not `/api/v3.1/documents`. `docxtract.rpatech.ai` is a legacy alias that redirects here. Do not use it. Endpoints have **no file extension**. The older `.php` forms — `/v3.1/documents.php` — still work and always will, so existing integrations are unaffected, but new code should use the extensionless paths above. ## Versions | Version | Status | |---|---| | v1, v2 | Retired — 404 | | v3 | Live, maintenance only. No new integrations. | | v3.1 | Current. This document. | ## Terminology `model` and `document_type` both refer to a **prompt template name** in the `prompt_templates` table (`invoice`, `passport`, `aadhaar`, …) — *not* an LLM name. The underlying AI provider and LLM are deliberately not exposed in production. ## Single-page vs multi-page `POST /v3.1/documents` behaves differently based on server-side page count: - **PDF of 3 pages or fewer, or any image** → processed synchronously, `200` with the result. - **PDF of more than 3 pages** → split into chunks, `202` with a manifest. The caller then calls `process` once per chunk and `result` to collect the stitched result. The threshold is `MP_ASYNC_THRESHOLD` (3). Callers cannot opt out, so **any integration that accepts arbitrary PDFs must handle the `202` branch.** Official SDKs do this transparently. ## Credits Billing is per page (1 page = 1 credit), not per request. Note that a failed extraction (`extraction_failed`, 422) still deducts 1 credit on the synchronous path. ## Result retention Results are **not stored** unless you pass `store_db: true`. Without it the extracted data exists only in the response, and no `extraction_id` is returned. contact: name: DocXtract url: https://docxtract.io x-built-by: RPATech servers: - url: https://api.docxtract.io description: Production security: - BearerAuth: [] tags: - name: Extraction description: Synchronous single-document extraction - name: Multi-page description: Split / process / collect flow for PDFs above the page threshold - name: Discovery description: Capability and health checks that do not consume credits paths: /v3.1/documents: post: tags: [Extraction] operationId: extractDocument summary: Extract structured data from a document description: | Upload a PDF or image and receive structured JSON. Returns `200` with the extracted result for images and PDFs of 3 pages or fewer. Returns `202` with a chunk manifest for PDFs above that threshold — see the Multi-page flow. requestBody: required: true content: multipart/form-data: schema: type: object required: [file] properties: file: type: string format: binary description: PDF, JPG, or PNG. Maximum 10 MB. Maximum 150 pages. options: type: string description: | JSON-encoded options object, sent as a form field (a string, not a nested part). See the `ExtractOptions` schema. example: '{"model":"invoice","store_db":true}' responses: "200": description: Extraction complete (synchronous path) headers: X-RateLimit-Limit: { $ref: "#/components/headers/XRateLimitLimit" } X-RateLimit-Remaining: { $ref: "#/components/headers/XRateLimitRemaining" } X-RateLimit-Reset: { $ref: "#/components/headers/XRateLimitReset" } content: application/json: schema: { $ref: "#/components/schemas/ExtractionSuccess" } "202": description: | Document was split for multi-page processing. `data` is the chunk manifest — call `process` once per chunk `job_id`, then `result` with the parent `job_id`. content: application/json: schema: { $ref: "#/components/schemas/SplitManifest" } "400": { $ref: "#/components/responses/BadRequest" } "401": { $ref: "#/components/responses/Unauthorized" } "402": { $ref: "#/components/responses/PaymentRequired" } "413": { $ref: "#/components/responses/PayloadTooLarge" } "422": { $ref: "#/components/responses/ExtractionFailed" } "429": { $ref: "#/components/responses/RateLimited" } "500": { $ref: "#/components/responses/ServerError" } "503": { $ref: "#/components/responses/ServiceUnavailable" } /v3.1/process: post: tags: [Multi-page] operationId: processChunk summary: Process one chunk of a split document description: | Call once per chunk `job_id` from the `202` manifest. Chunks may be processed sequentially or in parallel. The model is chosen **here**, per chunk — the split itself is model-independent. Passing different models across chunks of one document is permitted but produces a `mixed_models` warning on collect. Safe to retry: a chunk already `done` replays idempotently with code `chunk_already_processed` and is not billed twice. Credits are deducted per successful chunk, only after the result is persisted. parameters: - name: job_id in: query required: true schema: { type: string } description: Chunk job id (not the parent job id). May also be sent as a form field. requestBody: required: false content: multipart/form-data: schema: type: object properties: job_id: type: string description: Alternative to the query parameter. options: type: string description: | Same JSON-encoded options as `documents`. The form field must be named exactly `options`. example: '{"model":"invoice"}' responses: "200": description: | Chunk processed, or replayed if already done. `data` is always `null` — the extracted content is retrieved via `result`. content: application/json: schema: { $ref: "#/components/schemas/ChunkProcessed" } "400": description: Missing `job_id`, malformed `options`, or unknown model. content: application/json: schema: { $ref: "#/components/schemas/Error" } "401": { $ref: "#/components/responses/Unauthorized" } "402": { $ref: "#/components/responses/PaymentRequired" } "404": { $ref: "#/components/responses/JobNotFound" } "409": description: | Another request holds the single-flight claim on this chunk. Retry shortly; a stuck claim becomes re-claimable after 120 seconds. content: application/json: schema: { $ref: "#/components/schemas/Error" } "410": { $ref: "#/components/responses/JobGone" } "422": { $ref: "#/components/responses/ExtractionFailed" } "429": { $ref: "#/components/responses/RateLimited" } "500": { $ref: "#/components/responses/ServerError" } /v3.1/result: get: tags: [Multi-page] operationId: collectResult summary: Collect the stitched multi-page result description: | Streams every completed chunk through the JSON stitcher in page order and returns the combined result. A pure read by default — re-fetchable any number of times within the job TTL (2 hours from split). Returns `status: "partial"` with `pending_pages` and `failed_pages` when chunks are still outstanding, so it doubles as a progress poll. Passing `finalize=true` hard-deletes all extracted data for the job after the response is flushed. This is irreversible — do not finalize until the result is safely stored on the caller's side. parameters: - name: job_id in: query required: true schema: { type: string } description: Parent job id from the `202` manifest. - name: finalize in: query required: false schema: { type: boolean, default: false } description: | When true, permanently deletes all extracted data for this job after responding. Irreversible. responses: "200": description: Combined result. `status` is `complete` or `partial`. content: application/json: schema: { $ref: "#/components/schemas/CollectedResult" } "400": description: Missing `job_id`. content: application/json: schema: { $ref: "#/components/schemas/Error" } "401": { $ref: "#/components/responses/Unauthorized" } "404": { $ref: "#/components/responses/JobNotFound" } "410": { $ref: "#/components/responses/JobGone" } "429": { $ref: "#/components/responses/RateLimited" } "500": { $ref: "#/components/responses/ServerError" } /v3.1/models: get: tags: [Discovery] operationId: listModels summary: List document models available to this key description: | Returns the prompt templates (document types) the authenticated key may use. **Consumes no credits.** Authentication here is a direct key lookup rather than the standard `ApiAuth` path, specifically to avoid deduction — making this the correct endpoint for capability discovery and for populating UI dropdowns. Returns an empty `models` array — not an error — when the key has no configured model access. security: - BearerAuth: [] - ApiKeyHeader: [] - ApiKeyQuery: [] responses: "200": description: Allowed models for this key. content: application/json: schema: { $ref: "#/components/schemas/ModelList" } "401": description: Missing, invalid, inactive, or expired key. content: application/json: schema: { $ref: "#/components/schemas/Error" } "405": description: Method not allowed — use GET. content: application/json: schema: { $ref: "#/components/schemas/Error" } "500": { $ref: "#/components/responses/ServerError" } /v3.1/health: get: tags: [Discovery] operationId: healthCheck summary: Service health description: Unauthenticated liveness check. Reports API version and database connectivity. security: [] responses: "200": description: Service status. content: application/json: schema: { $ref: "#/components/schemas/Health" } /v3.1/authorised: get: tags: [Discovery] operationId: checkAuthorisation summary: Verify an API key is active description: | Lightweight key check — confirms a key is active without processing a document. Use this to validate a key at integration setup time. responses: "200": description: Key status. content: application/json: schema: type: object properties: success: { type: boolean } data: { type: object, additionalProperties: true } "401": { $ref: "#/components/responses/Unauthorized" } components: securitySchemes: BearerAuth: type: http scheme: bearer description: | `Authorization: Bearer sk_...` DocXtract keys are `sk_` followed by 32 hex characters. Note the **underscore** — A hyphen after `sk` means the key belongs to a different API provider, not DocXtract. ApiKeyHeader: type: apiKey in: header name: X-API-Key description: Accepted by `models` only. ApiKeyQuery: type: apiKey in: query name: api_key description: | Accepted by `models` only. Discouraged — keys in query strings leak into server logs, proxies, and browser history. Prefer the Bearer header. headers: XRateLimitLimit: description: Requests permitted in the current window. schema: { type: integer } XRateLimitRemaining: description: Requests remaining in the current window. schema: { type: integer } XRateLimitReset: description: Unix timestamp when the current window resets. schema: { type: integer } schemas: ExtractOptions: type: object description: | Sent as a **JSON-encoded string** in the `options` form field, not as structured multipart data. Only `model`, `store_db` and `document_type` are accepted. Any other key is silently ignored. properties: model: type: string description: | Prompt template name — the document type. Call `models` for the list available to your key. examples: [invoice, passport, aadhaar, bank_statement, salary_slip] document_type: type: string default: invoice description: | Synonym for `model`; either key works. Lowercased and trimmed server-side. **Defaults to `invoice`** when omitted — an unrelated document sent without this field is extracted as an invoice, so always set it explicitly. store_db: type: boolean default: false description: | Store the result server-side. **Defaults to false** — nothing is retained beyond the response unless you opt in. When true, the response carries an `extraction_id` you can reference later. ExtractionSuccess: type: object description: | Metadata sits at the **root level** alongside `data`, not nested under a `meta` key. required: [success, data] properties: success: { type: boolean, const: true } data: type: object additionalProperties: true description: | Extracted fields. **Shape varies by document type** — an invoice returns vendor/totals/line_items, a passport returns name/number/dob. Only the envelope is stable. Treat this as an open map, not a fixed schema. processing_time_ms: type: integer description: AI call duration when measurable, otherwise total wall-clock. model_used: type: string default: generic description: Prompt template actually used. pages: type: [integer, "null"] description: Page count discovered in the result payload. extraction_id: type: string description: | Stored result id. **Present only when `store_db` is true**, which is not the default. Named `extraction_id`, not `id`. SplitManifest: type: object description: The `202` response when a PDF exceeds the page threshold. required: [success, data, job_id] properties: success: { type: boolean, const: true } data: type: array description: Chunks to process, in page order. items: type: object properties: job_id: type: string description: Chunk job id — pass this to `process`. pages: type: string description: Page range covered, e.g. `"1-3"`. examples: ["1-3", "4-6"] job_id: type: string description: Parent job id — pass this to `result`. pages: type: integer description: Total pages in the document. result_url: type: string format: uri description: | Absolute URL of the collect endpoint for this job. Safe to follow directly. expires_at: type: string format: date-time description: Job expiry — 2 hours from split. Chunks are unrecoverable after this. message: type: string description: Human-readable next-step hint. ChunkProcessed: type: object required: [success, job_id, status, code] properties: success: { type: boolean, const: true } data: type: "null" description: Always null. Retrieve content via `result`. job_id: { type: string, description: Chunk job id. } status: { type: string, const: done } code: type: string enum: [chunk_processed, chunk_already_processed] description: | `chunk_already_processed` means this chunk was already done and the call replayed idempotently — no credits were charged. pages: { type: string, description: "Page range, e.g. \"4-6\"." } chunks_done: { type: integer } chunks_total: { type: integer } processing_time_ms: type: integer description: Present on fresh processing; absent on idempotent replay. CollectedResult: type: object required: [success, data, job_id, status] properties: success: { type: boolean, const: true } data: type: object additionalProperties: true description: Stitched result across all completed chunks, in page order. job_id: { type: string } status: type: string enum: [complete, partial] description: "`partial` means chunks are still pending or have failed." pages_total: { type: integer } pages_processed: { type: integer } chunks_done: { type: integer } chunks_total: { type: integer } credits_used: { type: integer } expires_at: { type: string, format: date-time } pending_pages: type: array items: { type: integer } description: Present only when chunks are outstanding. failed_pages: type: array items: { type: integer } description: Present only when chunks failed. Retry those chunk job_ids via `process`. warnings: type: array items: { type: string } description: | Present when chunks were processed with different models (`mixed_models`) — the combined result may be inconsistent. finalized: type: boolean description: Present and true when `finalize=true` was passed. Data has been deleted. ModelList: type: object properties: success: { type: boolean, const: true } data: type: object properties: models: type: array description: Empty when the key has no configured model access. items: type: object additionalProperties: true Health: type: object properties: status: { type: string, examples: [ok] } timestamp: { type: integer } version: { type: string, examples: ["v3.1"] } environment: { type: string, enum: [production, development] } database: { type: string, examples: [connected] } Error: type: object required: [success, error] properties: success: { type: boolean, const: false } error: type: object required: [code, message] properties: code: type: string description: Stable machine-readable code. Branch on this, never on `message`. enum: - invalid_api_key - expired_api_key - usage_limit_exceeded - insufficient_credits - rate_limit_exceeded - too_many_open_jobs - invalid_request - invalid_file - invalid_file_type - file_too_large - invalid_options - unknown_model - page_limit_exceeded - method_not_allowed - job_not_found - job_expired - chunk_in_progress - chunk_source_lost - extraction_failed - persist_failed - server_error - server_busy message: type: string description: Human-readable. Wording is not stable across releases. details: description: Present on some errors, e.g. `accepted_types`, `job_id`, `pages`. type: object additionalProperties: true responses: BadRequest: description: | Malformed request — `invalid_file`, `invalid_file_type`, `invalid_options`, `invalid_request`, `page_limit_exceeded` (over 150 pages), or `unknown_model`. content: application/json: schema: { $ref: "#/components/schemas/Error" } Unauthorized: description: "`invalid_api_key` or `expired_api_key`." content: application/json: schema: { $ref: "#/components/schemas/Error" } PaymentRequired: description: | `usage_limit_exceeded` — the key's `allowed_requests` cap is reached — or `insufficient_credits` when remaining credits do not cover the document's pages. content: application/json: schema: { $ref: "#/components/schemas/Error" } PayloadTooLarge: description: "`file_too_large` — over the 10 MB limit." content: application/json: schema: { $ref: "#/components/schemas/Error" } ExtractionFailed: description: | `extraction_failed` — the document was accepted but extraction did not succeed. **This still deducts 1 credit** on the synchronous path. On the multi-page path the chunk stays available for retry. content: application/json: schema: { $ref: "#/components/schemas/Error" } JobNotFound: description: | `job_not_found`. Jobs are key-scoped, so another key's job is indistinguishable from a nonexistent one. content: application/json: schema: { $ref: "#/components/schemas/Error" } JobGone: description: | `job_expired` (past the 2-hour TTL, or already finalized) or `chunk_source_lost` (staged source file no longer on disk). Both require re-uploading the document. content: application/json: schema: { $ref: "#/components/schemas/Error" } RateLimited: description: | `rate_limit_exceeded` — default limits are 10/minute, 100/hour, 1000/day, 10000/month — or `too_many_open_jobs` (maximum 5 unexpired multi-page jobs per key). headers: X-RateLimit-Limit: { $ref: "#/components/headers/XRateLimitLimit" } X-RateLimit-Remaining: { $ref: "#/components/headers/XRateLimitRemaining" } X-RateLimit-Reset: { $ref: "#/components/headers/XRateLimitReset" } content: application/json: schema: { $ref: "#/components/schemas/Error" } ServerError: description: "`server_error`, or `persist_failed` when extraction succeeded but storage did not (no credits charged — retry the chunk)." content: application/json: schema: { $ref: "#/components/schemas/Error" } ServiceUnavailable: description: "`server_busy` — staging capacity exhausted. Retry in a few minutes." content: application/json: schema: { $ref: "#/components/schemas/Error" }