jalapenojson

jalapenojson

A binary format for rows of data

A text header names each column with its type and its size in bytes. The values follow in binary, column by column, each at its column's size, so a reader can calculate where any value starts instead of parsing everything before it. There is a reader in C, and a reader and writer in Python and in JavaScript.

Format version 0.7, experimental

Compared with JSON

The same three rows in both formats. JSON repeats every key in every row and writes every value as text. jalapenojson writes the column names once, in a header of text, and stores the values in binary, one column after another. It is shown here the way the spec writes a document: the header as text, then each byte after it as two hex digits, with a comment naming the column and its values.

A document does not need to be readable in a plain text editor the way JSON is. The header can be read anywhere, and a plugin for your IDE or an editor of its own can show the values as rows.

JSON222 bytes minified
JSON
[
  {"pepper": "jalapeño", "scoville": 8000, "picked": "2026-09-14", "ripe": false},
  {"pepper": "serrano", "scoville": 23000, "picked": "2026-09-21", "ripe": false},
  {"pepper": "habanero", "scoville": 350000, "picked": "2026-10-02", "ripe": true}
]
jalapenojson123 bytes
jalapenojson
0.7
#pepper:s:9 #scoville:i:4 #picked:d:4 #ripe:y:1
3
00 00                                padding
6a 61 6c 61 70 65 c3 b1 6f           pepper: jalapeño
73 65 72 72 61 6e 6f 00 00           pepper: serrano
68 61 62 61 6e 65 72 6f 00           pepper: habanero
00 00 00 00 00                       padding
40 1f 00 00 d8 59 00 00 30 57 05 00  scoville: 8000, 23000, 350000
00 00 00 00                          padding
e6 50 00 00 ed 50 00 00 f8 50 00 00  picked: 2026-09-14, 2026-09-21, 2026-10-02
00 00 00 00                          padding
00 00 01                             ripe: false, false, true

How a reader finds a value

The same document again, one box per byte: the header, then each column. Every value in a column has the same size, so the position of any value can be calculated from where its column starts. Hover over a field, or move to it with Tab, to see the calculation.

This view needs JavaScript. The document above has the same bytes, written in hex.

How it works

The header

The header is text. It lists each column as #name:type:size: #scoville:i:4 is an integer column whose values take 4 bytes each. Its last line counts the rows. Everything after the header is binary and holds only the values. There are no keys or separators and nothing is escaped, so a text value can contain a newline or a quote.

Fixed sizes

Every value in a column takes up exactly the column's size. Text is UTF-8, and a value shorter than its column is followed by a NUL, the break byte, after which the reader ignores the rest of the field. A size counts bytes, so jalapeño needs 9 because ñ is 2 bytes in UTF-8. The other types are little-endian binary at a width the type sets: an i is 1, 2, 4 or 8 bytes and a d is 4.

Finding a value

The values are stored column by column, and each column starts at a multiple of 8 from the first byte. Row N of a column starts at the column's start plus N times its size. The header's sizes and counts give every column's start, so the reader goes to that one field and reads nothing else. It needs no index, and a column of numbers can be used in place as an array.

A column can also hold a list of rows that use another schema. Each list column has a table of its own after the document's rows, stored by column the same way, and its field is 16 bytes: the row where the list starts in that table, then how many rows it holds. So a value inside a list is found by the same arithmetic. Lists can be nested to any depth, and the format has no limit on the number of rows or columns or on column sizes.

Column types

  • i integer, 1, 2, 4 or 8 bytes
  • f float, 8 bytes
  • s text, any size
  • b raw bytes, any size
  • y boolean, 1 byte
  • d date, 4 bytes
  • t time of day, 8 bytes
  • n date and time, 8 bytes
  • z date and time with offset, 10 bytes
  • 2 list of rows using schema 2, 16 bytes

When to use it

How it compares with JSON, based on the benchmark and on how the format works.

Good fit

  • Reading part of a large document, such as one row or a few columns. The reader reads only those values, while a JSON parser has to parse the whole document first. This is where the difference is largest, in every language, and a column of numbers is already an array in the document's bytes.
  • Reading or writing a whole table in JavaScript or C, where reading is faster than JSON.parse and yyjson, and writing is faster than JSON.stringify.
  • Data stored or sent uncompressed, such as a cache, a queue or a file on disk. Column names are written once, numbers and dates are binary, and values have no quotes or separators, so it is about half the size of the same JSON.
  • Nested data, such as orders with their line items. A list column keeps the size saving without flattening the data into one row per item.

Small difference

  • Sending it gzipped or brotli-compressed. Compression removes much of what JSON repeats, so it is smaller and faster by less than it is raw.
  • A small document loaded once by a web page. Most of the time goes to the download, and the JavaScript reader's first call to read every value is slower than JSON.parse, so the two end up about even.
  • Reading every row into dicts in Python. It is faster than json.loads, but by less than in the other languages, and by little on varied data.

Poor fit

  • Free text. Every value in a column takes up the size of the longest one, so text that varies a lot in length wastes space.
  • Rows with different fields. Every row has every column, and there is no null.
  • Files edited by hand. The values are binary, so a plain text editor does not show them, and a value cannot grow past its column's size without rewriting the column and every column after it.

Benchmark results

Reading and writing the same rows with JSON and with jalapenojson in each language. Every result is checked against a checksum of the rows. There are two data sets, and they are shown separately: repetitive rows, similar to typical API output, and high-entropy rows with random values that do not compress well.

Nested orders carry one to six line items each.

50,000 rowsJSONjalapenojsondifference
repetitive, raw5.55 MB2.40 MB57% smaller
repetitive, gzipped609 KB430 KB29% smaller
high-entropy, raw6.27 MB3.00 MB52% smaller
high-entropy, gzipped2.27 MB1.90 MB16% smaller
nested orders, raw12.04 MB5.49 MB54% smaller
nested orders, gzipped1.78 MB1.43 MB20% smaller

Measured on 2026-10-07 at commit 5548d6a: Intel Xeon Processor @ 2.80GHz, 4 cores, Linux 6.18.44-fc-v77; Node 22.22.0, Python 3.13.16, cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 with -O2, cJSON 1.7.19, yyjson 0.13.0. The full report also has results for 1,000 rows, the first call in a new process, brotli and a slow network, and describes how each number was measured.

Usage

Python
from datetime import date, datetime, timezone, timedelta
from jalapenojson import encode, decode

blob = encode(rows, schema=[("id","i",4), ("name","s",16), ("seen","z")])
rows = decode(blob)
rows[0]["seen"]           # datetime(2026, 1, 15, 9, 30, tzinfo=timezone(timedelta(hours=2)))

The guide has more, including dates, booleans, nested lists, calculating column sizes and the rest of the C API.

Download

Each implementation is a single file with no dependencies. Copy it into your project.