Skip to content

Convert an email to JSON

An .eml file is text, but it is not data. It is MIME: nested parts, transfer encodings, a charset per part, encoded words in the subject, and — in anything that is a reply — a quoted copy of the whole previous conversation underneath the two lines that matter. This page shows what the same message looks like after all of that has been dealt with.

The output below is real. It is what MailMint’s parser (packages/parser, the code behind POST /v1/parse) returned on 25 September 2026 for fx-08-reply-chain-invoice.eml, a message from its test corpus, sent without a schema. It is printed in full except for one key, noted where it was cut.

What “email to JSON” has to mean #

Splitting a message into headers and body is the easy tenth of the job. What makes the result usable is everything a MIME library leaves to you:

  • Decoded text. Quoted-printable and base64 undone, legacy charsets such as ISO-8859-1 and Windows-1252 normalised to UTF-8, RFC 2047 encoded words in headers and RFC 2231 file names reassembled.
  • The part you actually want. A reply carries the whole thread. stripped_text is the new message with the quoted chain and the signature removed.
  • Text from HTML-only mail. When there is no plain-text part, text_from_html is rendered from the HTML so there is always something to read.
  • Tables as rows. HTML grids, repeating blocks and whitespace-aligned text all become tables[] with headers, rows, records and a row_count.
  • What is in it. Amounts with currency, dates normalised to ISO 8601, ids, email addresses, URLs, phone numbers, postal addresses and a guessed document type, under detected.

The request #

Send the raw message as raw_mime (plain text or base64, up to 25 MB). Leave schema out and you get the cleaned object without any extracted fields. Nothing is stored.

# jq builds the JSON body so the message does not need escaping by hand
jq -Rs '{raw_mime: .}' message.eml |
curl -X POST "$MAILMINT_URL/v1/parse" \
  -H "Authorization: Bearer $MAILMINT_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @-

The message itself, abbreviated to the parts that matter (the transport headers and a DKIM signature are above it in the file):

From: Tomas Berglund <tomas@halvard.se>
Subject: Re: Wrong total on INV-77211
Date: Mon, 24 Aug 2026 17:22:38 +0200
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

Sorted - the corrected invoice is INV-77213 and the total is $2,415.60 inc=
luding tax.
Due date is September 12, 2026.

Best,
Tomas

--
Tomas Berglund | Finance
Halvard AB | +46 8 121 489 00

On Mon, Aug 24, 2026 at 4:51 PM Florian Standhartinger <k7m2xq4h9bwz@parse=
.example.com> wrote:
> Hi Tomas,
>
> The invoice you sent (INV-77211) has last month's total on it - $1,980.0=
0.
> Could you reissue?
> ...

The response, in full #

{
  "received_at": "2026-09-25T23:35:02.006Z",
  "headers": {
    "message_id": "<CAF9y@mail.halvard.se>",
    "date": "2026-08-24T15:22:38.000Z",
    "subject": "Re: Wrong total on INV-77211",
    "from": { "name": "Tomas Berglund", "email": "tomas@halvard.se" },
    "to": [ { "name": null, "email": "k7m2xq4h9bwz@parse.example.com" } ],
    "cc": [],
    "reply_to": [],
    "in_reply_to": "<CAF3x@mail.example.com>",
    "references": [ "<CAF1x@mail.halvard.se>", "<CAF3x@mail.example.com>" ],
    "raw": { … every header exactly as received, cut here for length … }
  },
  "body": {
    "text": "Sorted - the corrected invoice is INV-77213 and the total is $2,415.60 including tax.\nDue date is September 12, 2026.\n\nBest,\nTomas\n\n--\nTomas Berglund | Finance\nHalvard AB | +46 8 121 489 00\n\nOn Mon, Aug 24, 2026 at 4:51 PM Florian Standhartinger <k7m2xq4h9bwz@parse.example.com> wrote:\n> Hi Tomas,\n>\n> The invoice you sent (INV-77211) has last month's total on it - $1,980.00.\n> Could you reissue?\n>\n> Thanks\n>\n> On Mon, Aug 24, 2026 at 2:10 PM Tomas Berglund <tomas@halvard.se> wrote:\n>> Invoice INV-77211 attached, total $1,980.00, due September 5, 2026.\n>>\n>> --\n>> Tomas",
    "html": null,
    "text_from_html": null,
    "stripped_text": "Sorted - the corrected invoice is INV-77213 and the total is $2,415.60 including tax.\nDue date is September 12, 2026.\n\nBest,\nTomas",
    "language": "en"
  },
  "attachments": [],
  "tables": [],
  "detected": {
    "type": "invoice",
    "emails": [],
    "urls": [],
    "phones": [],
    "amounts": [ { "value": 2415.6, "currency": "USD", "raw": "$2,415.60" } ],
    "dates": [ { "value": "2026-09-12", "raw": "September 12, 2026" } ],
    "ids": [
      { "kind": "invoice_number", "value": "INV-77213" },
      { "kind": "invoice_number", "value": "INV-77211" }
    ],
    "addresses": []
  },
  "fields": {},
  "flags": [ "no_schema" ],
  "needs_review": false,
  "parse": {
    "request_id": "req_local",
    "schema_version": null,
    "model": null,
    "llm_used": false,
    "timings_ms": { "total": 102, "mime": 49, "deterministic": 43, "llm": 0, "persist": 0 },
    "warnings": []
  }
}

Through the HTTP endpoint the same object also carries id, mailbox and raw_url — all null, because /v1/parse stores nothing — and an auth block, described below. The timings are from a local run on a busy machine; they are not a latency claim.

Key by key #

KeyWhat to use it for
headersDecoded, with addresses split into name and email and the date in UTC. in_reply_to and references let you thread conversations. raw keeps every header verbatim for anything not lifted out.
body.textThe plain-text part, transfer encoding undone — note the soft line breaks (inc=) are gone.
body.stripped_textThe new content only. Here the signature block and two levels of quoted reply are removed, which is what you want to show a person or hand to a model.
body.languageA detected language code for routing.
attachments[]One object per file with filename, content_type, size, sha256, inline and content_id. Empty here.
tables[]Every table found, with row_count and truncated. Empty here: a sentence is not a table.
detectedEverything recognisable, from the whole text. That is why both invoice numbers appear: INV-77211 only exists in the quoted part. detected is a list of candidates, not an answer — deciding which one you mean is what a schema is for.
flagsno_schema says, in the payload, that no fields were asked for — so an empty fields cannot be mistaken for a failed extraction.

Sender authentication #

For mail that arrives at a MailMint address, auth carries SPF, DKIM and DMARC verdicts computed by MailMint’s own SMTP server against live DNS. Through /v1/parse the situation is different, and the response says so rather than guessing: DKIM can be verified from the message and DNS alone, but SPF and DMARC need the connecting IP and the SMTP envelope, which do not exist when someone posts bytes to an HTTP endpoint. Those two come back as "unavailable" with a reason, not as a pass. The details.

Adding a schema #

Add a schema and the same request also returns fields: one entry per field with a value, a computed confidence, the source layer and the evidence it was read from. For this message you would ask for the invoice number and total, and the parser would have to choose between the corrected invoice and the one quoted below it — which is the kind of decision the confidence and evidence exist to make checkable.

"schema": [
  { "name": "invoice_number", "type": "string", "description": "the corrected invoice number" },
  { "name": "total",          "type": "currency", "description": "total including tax" },
  { "name": "due_date",       "type": "date" }
]

See parsing invoice emails for a full worked example with a field that came back doubtful, and the schema reference for the 13 field types.

Limits #

  • Up to 25 MB per message.
  • Every call counts as one parsed email against your plan — 300 a month on Free, no card.
  • Nothing is stored: no message row, no raw bytes, no attachment blobs. Attachment urls are null for that reason; use a mailbox if you need the bytes later.
  • Text inside PDF attachments is not extracted yet. The file’s metadata and checksum are returned, its contents are not.

When a library is enough #

If you only need headers and a body out of well-formed mail, in-process, a MIME library is free and has no network hop: Python’s standard email package, or mailparser on npm. The case for an API is everything after the split — the reply chain, the charsets real senders get wrong, tables, detection, and a schema with a confidence per field — and not having to maintain it.

Get a free API key or start with the quickstart.