The Synthesize Bio MCP Server enables AI agents to compare gene expression between sample groups, analyze bulk and single-cell RNA sequencing data, and run long-running bioinformatics workflows that return completed analysis reports.
Encrypted at rest, isolated from the model
Resolved from an AES-256-GCM vault at the moment of the call and attached to the request — the model never sees the secrets.
Try asking
Returns the structured-metadata schema used to turn a natural-language experiment description into sample groups. Response fields include `group_schema` (the JSON shape of one sample group), the `fields` and `perturbation_fields` inventory for the active metadata version, and `instructions` for building the `groups` array. The resulting `groups` array is the input to resolve_sample_metadata. This tool makes no AI calls and consumes no usage budget.
Deterministically harmonizes already-structured sample `groups` to Synthesize Bio's controlled ontology vocabulary. Does not accept a natural-language prompt and makes no AI calls. Input: `groups` — a JSON array of sample-group objects matching get_metadata_schema's `group_schema`. Returns a resolution_id and a per-group breakdown including tissue, disease, cell type/line (with resolved ontology ids), sex, age, and full perturbation details (type, label/id, gene mechanism and mechanism type, dose, dose count, timepoint). Each resolved field also carries ranked candidate options for disambiguation. If a value resolves to the wrong id, correcting the term in `groups` and calling again produces a new resolution. The `warnings` array flags issues such as a described drug that did not match the ontology. When status is `resolving`, the same resolution_id identifies the pending request and can be passed again to poll. A completed resolution_id is required by analyze_gene_expression. Optional workspace_id selects which workspace the resolution is stored under. When omitted, a single-workspace account is assigned automatically. When the account has more than one workspace and workspace_id is omitted, the call is rejected with failure_kind `workspace_selection_required` and a `workspaces` list of `{ name, workspace_id }`; no metadata resolution runs until the user picks and the tool is retried with workspace_id. Quota and monthly-limit errors include a request-higher-limits URL and indicate remaining samples/cells when available.
Starts a differential gene expression analysis using Synthesize Bio's AI platform. Requires the resolution_id returned by resolve_sample_metadata; raw natural-language prompts are not accepted. Requires `user_confirmed_metadata: true`. When the flag is missing or false, the call is rejected with failure_kind `user_confirmation_required`. Optional workspace_id selects which workspace owns the generated dataset. When omitted, the workspace from resolve_sample_metadata is used automatically. When the account has more than one workspace and the resolution has no workspace, the call is rejected with failure_kind `workspace_selection_required` and a `workspaces` list of `{ name, workspace_id }`. Returns a job_id immediately; get_analysis_results accepts that job_id and returns analysis status or results. The pipeline runs two steps: (1) GEM-1 — Synthesize Bio's Gene Expression Model inference; (2) Differential expression — GPU-accelerated DESeq2 (negative-binomial GLM with Wald test, Cook's outlier filter, and Benjamini-Hochberg padj). All genes are tested; pre-filtering is handled by DESeq2's independent filtering. If the query is unsupported, later polling responses include failure_kind `unsupported_query` and suggested_queries. Quota and monthly-limit errors include a request-higher-limits URL; a previous successful resolve does not grant an extra run when the account is out of budget.
Polls the status of a gene expression analysis. Each call waits server-side for a short bounded window and may return earlier if progress is detected. Responses always include a `structuredContent` object (declared by the tool's `outputSchema`); MCP clients read from `structuredContent` directly rather than re-parsing JSON out of the human-readable text. `structuredContent` always has `status` (one of `running`, `complete`, `failed`), `job_id`, and `steps_completed`. While running, it also has `step` (`gem_model` or `diff_expr`), `message`, and `progress_label`/`progress_percent`/`progress_bar`. Failed responses include `error`, and may also include `failure_kind`, `user_action_required`, and `suggested_queries`. When `status` is `complete`, `structuredContent` carries: `metadata` (prompt, modality, groups, plus summary counts such as `significant_genes`, `significant_up`, `significant_down`, `total_genes_tested`); `results` — up to 1000 differential expression rows (each with `gene_id`, `gene_symbol`, `log2FoldChange`, `pvalue`, `padj`, `neg_log10_padj` (pre-computed `-log10(padj)`, clamped to 300 if padj underflows), `direction`, `significant`) suitable for downstream analysis or visualization (e.g. a volcano plot with x = `log2FoldChange`, y = `neg_log10_padj`); `plot_results` — the top ~200 most significant rows (same per-row shape, pre-sorted most-significant-first), pre-sliced for charting; the full `results` array is better suited to tables, summaries, and analysis; `results_returned` and `results_total` for truncation accounting; `plot_results_returned` for the plotted subset size; `dataset_link` — `{ dataset_id, title, url }` for the Synthesize Bio platform dataset (or `null`). The accompanying `content[0].text` is a human-readable Markdown summary of the same data. For hosts that do not surface `structuredContent` (e.g. claude.ai), it inlines only the top ~200 most significant rows as an array of objects under a top-level `results` key — same per-row schema as `structuredContent.results`, including the pre-computed `neg_log10_padj` field — so chart-widget code can use those rows directly. The full result set remains available via `structuredContent.results` and the dataset link when present.
Returns a presigned S3 URL to download the raw gene expression counts data (JSON) produced by a completed or in-progress analysis job. Supported only in Claude Code or Codex with control of local terminal sessions; chat clients and hosted/browser-based code sandboxes are unsupported, even with Python or shell tools. The data is typically large (20,000+ genes by N samples) and is intended for local processing. Raw counts are optional for custom computations; DE results and contrast labels are available separately from get_analysis_results. The JSON has the following structure: { gene_order: string[] (Ensembl IDs), outputs: [{ counts: number[], metadata: object }], model_version: string }. Each entry in 'outputs' corresponds to one sample; 'counts' is aligned with 'gene_order'. The response also returns a second presigned URL to a small (~500 KB) Parquet file mapping every Ensembl `gene_id` in `gene_order` to its HGNC gene_name — both files join on gene_id to label genes by symbol. Both URLs expire after 1 hour; fresh responses contain fresh URLs. Available after analyze_gene_expression has completed the GEM-1 step.
One endpoint, the same key, whichever client you use.
~/Library/Application Support/Claude/claude_desktop_config.json (Mac) · %APPDATA%\Claude\claude_desktop_config.json (Windows)
Replace API_KEY with your own key.
Already have an "mcpServers" section in your config? Just add the server entry inside it.
Discovery, routing, credentials, tool scoping and execution logs all happen at the gateway→connections stay ACTIVE with no work from you
Synthesize Bio MCP runs through a gateway that holds the credentials, scopes the access and records every call.
Managed auth, hosted MCP servers, and every Gmail tool your agent needs.
Free to start.