Gemina automates business document processing with AI-powered OCR, structured data extraction, and document search. Extract invoice headers and line items, define custom extraction templates for contracts and forms, tag files with metadata, and query indexed documents to calculate totals and aggregate records by vendor, currency, or date.
Encrypted at rest, isolated from the model
Resolved from an AES-256-GCM vault at the moment of the call and attached to the request — the model never sees the secrets.
Try asking
Reserve a pre-signed PUT slot for a forthcoming file upload. Returns the upload URL, an ``upload.headers`` dict (the agent MUST echo every header in this dict on the PUT -- today that is just ``Content-Type``), the slot expiry (5 minutes), and a ``next_tool_call`` recipe -- ``tag_file`` by default, or ``extract_document`` when ``purpose='extract'`` -- copy-paste the ``file_id`` into that follow-up call. This is the canonical path for any file the agent holds locally; bytes never traverse the LLM context (they go directly from the agent host to GCS). Allowed types: PDF, PNG, JPEG, GIF, WebP, HEIC/HEIF, AVIF. Max size 50 MB. The returned ``file_id`` (``gfile_…``) is an upload-slot handle accepted by both ``tag_file`` and ``extract_document`` (one slot type; each slot is single-use). It is NOT a document id.
Run the FileTag pipeline against a previously uploaded slot. The ``file_id`` comes from a prior ``files_create_upload`` call. The server validates the uploaded blob (size, content-type, optional SHA-256), atomically consumes the slot, runs the FileTag extraction (renaming + metadata embedding), and returns the structured result with the extracted metadata, the suggested filename, the ``enriched_file_url`` (short-lived signed URL to the renamed copy with metadata embedded into document properties), and a ``next_action`` recipe (``http_get_and_save``) telling the agent to download that URL and save it as the suggested filename -- act on it unless the user explicitly asked for metadata only. Each slot is single-use; reserve a new slot with ``files_create_upload`` to retry. The returned ``document_id`` is the handle for ``add_document_extractions``: run paid Core-OCR types (invoice_headers, invoice_line_items, ...) on this same stored document later without re-uploading.
Fetch a remote URL server-side and run the FileTag pipeline. The bytes never traverse the LLM context -- the agent supplies the URL, the server fetches under strict SSRF guards (HTTPS only, no private IP ranges, 30-second timeout, 50 MB cap, redirects disabled), and returns the structured tag result with metadata, suggested filename, ``enriched_file_url`` (short-lived signed URL to the renamed copy with metadata embedded into document properties), and a ``next_action`` recipe (``http_get_and_save``) telling the agent to download that URL and save it as the suggested filename -- act on it unless the user explicitly asked for metadata only. Use this when the file already lives at a public URL.
Search your indexed documents; no per-query or per-page charge (included with a Document Intelligence plan, fair-use rate-limited). Ask across the whole collection instead of fetching and parsing files yourself. Modes: 'structured' (exact filters over extracted fields: vendorName, docNumber, documentType, currency, issueDateFrom/To, totalAmountMin/Max, endUserId, ...), 'semantic' (natural-language similarity over extracted fields and FileTag metadata/summaries — not raw document body text), 'hybrid' (keyword + semantic fused with Reciprocal Rank Fusion — best default for free-text questions). Returns matched documents with their extracted fields; semantic and hybrid modes also return relevance scores. Documents you tag (FileTag) or run a structured extraction on are indexed automatically once document indexing is enabled for the tenant (OFF by default — enable it in account settings); plain 'ocr' extractions are not indexed. An empty result may mean indexing is not enabled, or that nothing matched the filters. (For a plain chronological list of your own extraction requests use ``list_extractions`` instead.)
Spend analytics over your indexed documents; no per-query or per-page charge (included with a Document Intelligence plan, fair-use rate-limited). Answer money questions across the whole collection instead of exporting and adding up files yourself. Compute sums/averages/min/max/counts, optionally grouped (vendor_name, currency, document_type, expense_type, payment_method, end_user_id, month, year) and filtered (same filters as query_documents). Example: total spent per vendor in Q3 = metrics [{'op':'sum','field':'total_amount'}], group_by ['vendor_name'], filters {issueDateFrom, issueDateTo}. Counts are per DOCUMENT: a document's multiple extractions collapse to one, and re-uploads collapse when their full invoice identity (vendor_tax_id, doc_number, issue_date) matches. NOTE: money metrics are always split per currency unless a currency filter is given (mixed-currency totals would be meaningless); the response meta flags when that grouping was added automatically. Documents you tag or run a structured extraction on are indexed once the tenant enables document indexing (OFF by default); plain 'ocr' and skipped (fieldless / uncredited) extractions do not appear.
(Re)index one of the tenant's documents into the searchable index — use after corrections, or to backfill a document processed before indexing was enabled. Indexing normally happens automatically on every extraction once the tenant enables document indexing; this tool is the manual trigger. Returns per-outcome counts (indexed / skipped_opt_out / skipped_state / skipped_no_fields / skipped_not_in_plan / skipped_no_credits).
Run Core-OCR extraction on a previously uploaded file slot. The ``file_id`` comes from a prior ``files_create_upload`` call (``purpose='extract'`` yields the matching recipe; any slot works). Upload the bytes to the returned signed URL first; each slot is single-use. To extract more types from a document that is already stored, call ``add_document_extractions`` with its ``document_id`` instead of re-uploading. Choose one or more ``extraction_types``: 'ocr' (full text), 'invoice_headers', 'invoice_line_items', 'document_details_hebrew', 'document_line_items_hebrew', or 'custom_template' (requires a READY ``template_id``). Extraction is asynchronous: the call returns within seconds with either the completed result (fast documents) or an IN_PROCESS status carrying ``meta.correlationId`` — poll ``get_extraction_result`` with that id until complete. Duplicate protection is opt-in: pass an ``external_id`` of your own to enable it (re-submitting the same external_id within the dedup window idempotently returns the prior result, or errors if the file or extraction types differ; ``allow_duplicate=true`` overrides). Without an external_id every extraction is billed as new. Optional advanced knobs mirror the REST API: ``model_type`` selects the extraction model (see its recommendation per extraction type), and ``thinking`` / ``evaluation`` / ``correction`` / ``include_coordinates`` toggle accuracy passes and coordinate output.
Poll for the result of an asynchronous ``extract_document`` call. Pass the ``meta.correlationId`` from the extract response. Returns the completed extraction result once processing finishes, or an IN_PROCESS status while it is still running — poll again after a few seconds. Only correlations created by the calling API key are visible.
List the tenant's past document extractions, newest first. Filter by ``external_id`` (the idempotency key passed to extract_document), ``end_user_id``, and/or an ISO date window (``from_date``/``to_date``, YYYY-MM-DD). Paginate with ``skip``/``limit``. Returns extraction summaries — fetch full extracted data for one item with ``get_extraction``. This is the chronological log of your own extraction requests (a billing/audit view). To find documents by their content or extracted fields use ``query_documents``; to check on a call that is still running poll ``get_extraction_result`` with its correlationId rather than searching this list.
Fetch one extraction by its id (from list_extractions or a completed extract_document result), including the full extracted data. Only extractions created by the calling API key are visible.
Fetch one document by its id, including all of its extractions. Only documents created by the calling API key are visible.
Run additional extraction types on a document Gemina already stores — no re-upload — either to read the extracted values now or to make the document searchable for later. Pass the ``document_id`` (from ``tag_file``, ``extract_document`` or ``get_document``) and one or more ``extraction_types``: 'ocr', 'invoice_headers', 'invoice_line_items', 'document_details_hebrew', 'document_line_items_hebrew', or 'custom_template' (requires a READY ``template_id``). The stored OCR/image artifacts are reused; only the extraction step runs. Every structured type (all of the above except plain 'ocr') also adds this document's fields to your searchable collection — a richer field set than tagging alone indexes — so you can later find and total it with ``query_documents`` / ``aggregate_documents`` instead of re-processing the file. CHOOSE HOW TO WAIT. ``wait=true`` (the default) holds the call a few seconds so a fast document comes back already complete; if the status is still IN_PROCESS, poll ``get_extraction_result`` with ``pollCorrelationId``. ``wait=false`` returns immediately — use it when you are filing the document for future search and do not need the values in this turn; extraction and indexing still run to completion server-side, and the outcome (including failures) shows up afterwards in ``get_document`` / ``list_extractions``, so nothing is lost by not polling. Billed per extraction like an upload whichever mode you choose (paid credits — FileTag's free allowance does not cover it; a free-tier account gets CREDIT_EXHAUSTED with an upgrade link). Indexing additionally requires a Document Intelligence plan and the account's document-indexing setting switched on. Only documents created by the calling API key are eligible. Rejected (409) when a successful extraction of the same type already exists (use ``get_extraction``), when the document is still processing, or when it has recorded errors; 410 when its content was purged; 422 when the document has more pages than the requested type allows. Returns the whole document (all extractions, old and new) plus ``newExtractionIds`` and ``pollCorrelationId``.
Submit verified/corrected field values for a completed extraction — the extraction-quality feedback loop. ``data`` keys use the ``label:<human label>|ptr:/<json pointer>`` format addressing fields of the extraction result, e.g. {"label:Total Amount|ptr:/totalAmount": "118.00", "label:Vendor Name|ptr:/vendorName": "ACME Ltd"}. Returns a per-field comparison summary (correct/incorrect/missing counts). Each extraction accepts feedback once.
One endpoint, the same key, whichever client you use.
~/Library/Application Support/Claude/claude_desktop_config.json (Mac) · %APPDATA%\Claude\claude_desktop_config.json (Windows)
Replace API_KEY with your own key.
Already have an "mcpServers" section in your config? Just add the server entry inside it.
Discovery, routing, credentials, tool scoping and execution logs all happen at the gateway→connections stay ACTIVE with no work from you
Gemina MCP runs through a gateway that holds the credentials, scopes the access and records every call.
Managed auth, hosted MCP servers, and every Gmail tool your agent needs.
Free to start.