A text header names each column with its type and its size in bytes. The values follow in binary, column by column, each at its column's size, so a reader can calculate where any value starts instead of parsing everything before it. There is a reader in C, and a reader and writer in Python and in JavaScript.
The same three rows in both formats. JSON repeats every key in every row and writes every value as text. jalapenojson writes the column names once, in a header of text, and stores the values in binary, one column after another. It is shown here the way the spec writes a document: the header as text, then each byte after it as two hex digits, with a comment naming the column and its values.
A document does not need to be readable in a plain text editor the way JSON is. The header can be read anywhere, and a plugin for your IDE or an editor of its own can show the values as rows.
The same document again, one box per byte: the header, then each column. Every value in a column has the same size, so the position of any value can be calculated from where its column starts. Hover over a field, or move to it with Tab, to see the calculation.
This view needs JavaScript. The document above has the same bytes, written in hex.
しくみ
How it works
The header
The header is text. It lists each column as #name:type:size: #scoville:i:4 is an integer column whose values take 4 bytes each. Its last line counts the rows. Everything after the header is binary and holds only the values. There are no keys or separators and nothing is escaped, so a text value can contain a newline or a quote.
Fixed sizes
Every value in a column takes up exactly the column's size. Text is UTF-8, and a value shorter than its column is followed by a NUL, the break byte, after which the reader ignores the rest of the field. A size counts bytes, so jalapeño needs 9 because ñ is 2 bytes in UTF-8. The other types are little-endian binary at a width the type sets: an i is 1, 2, 4 or 8 bytes and a d is 4.
Finding a value
The values are stored column by column, and each column starts at a multiple of 8 from the first byte. Row N of a column starts at the column's start plus N times its size. The header's sizes and counts give every column's start, so the reader goes to that one field and reads nothing else. It needs no index, and a column of numbers can be used in place as an array.
A column can also hold a list of rows that use another schema. Each list column has a table of its own after the document's rows, stored by column the same way, and its field is 16 bytes: the row where the list starts in that table, then how many rows it holds. So a value inside a list is found by the same arithmetic. Lists can be nested to any depth, and the format has no limit on the number of rows or columns or on column sizes.
Column types
i integer, 1, 2, 4 or 8 bytes
f float, 8 bytes
s text, any size
b raw bytes, any size
y boolean, 1 byte
d date, 4 bytes
t time of day, 8 bytes
n date and time, 8 bytes
z date and time with offset, 10 bytes
2 list of rows using schema 2, 16 bytes
からさ
When to use it
How it compares with JSON, based on the benchmark and on how the format works.
Good fit
Reading part of a large document, such as one row or a few columns. The reader reads only those values, while a JSON parser has to parse the whole document first. This is where the difference is largest, in every language, and a column of numbers is already an array in the document's bytes.
Reading or writing a whole table in JavaScript or C, where reading is faster than JSON.parse and yyjson, and writing is faster than JSON.stringify.
Data stored or sent uncompressed, such as a cache, a queue or a file on disk. Column names are written once, numbers and dates are binary, and values have no quotes or separators, so it is about half the size of the same JSON.
Nested data, such as orders with their line items. A list column keeps the size saving without flattening the data into one row per item.
Small difference
Sending it gzipped or brotli-compressed. Compression removes much of what JSON repeats, so it is smaller and faster by less than it is raw.
A small document loaded once by a web page. Most of the time goes to the download, and the JavaScript reader's first call to read every value is slower than JSON.parse, so the two end up about even.
Reading every row into dicts in Python. It is faster than json.loads, but by less than in the other languages, and by little on varied data.
Poor fit
Free text. Every value in a column takes up the size of the longest one, so text that varies a lot in length wastes space.
Rows with different fields. Every row has every column, and there is no null.
Files edited by hand. The values are binary, so a plain text editor does not show them, and a value cannot grow past its column's size without rewriting the column and every column after it.
すうじ
Benchmark results
Reading and writing the same rows with JSON and with jalapenojson in each language. Every result is checked against a checksum of the rows. There are two data sets, and they are shown separately: repetitive rows, similar to typical API output, and high-entropy rows with random values that do not compress well.
Nested orders carry one to six line items each.
50,000 rows
JSON
jalapenojson
difference
repetitive, raw
5.55 MB
2.40 MB
57% smaller
repetitive, gzipped
609 KB
430 KB
29% smaller
high-entropy, raw
6.27 MB
3.00 MB
52% smaller
high-entropy, gzipped
2.27 MB
1.90 MB
16% smaller
nested orders, raw
12.04 MB
5.49 MB
54% smaller
nested orders, gzipped
1.78 MB
1.43 MB
20% smaller
Warmed up: JSON.parse and JSON.stringify against decode(), view() and encode(). One number column is read with numbers(), which hands it back as a Float64Array.
50,000 rows
data
JSON
jalapenojson
difference
read every value
repetitive
50.5 ms
13.3 ms
3.8x faster
high-entropy
45.7 ms
12.8 ms
3.6x faster
read every value, as columns
repetitive
60.7 ms
11.3 ms
5.4x faster
high-entropy
65.9 ms
11.1 ms
5.9x faster
read one number column
repetitive
50.2 ms
0.058 ms
872x faster
high-entropy
52.3 ms
0.058 ms
902x faster
read one text column
repetitive
49.1 ms
4.71 ms
10x faster
high-entropy
53.3 ms
2.39 ms
22x faster
read one row
repetitive
47.8 ms
0.0039 ms
12,299x faster
high-entropy
49.8 ms
0.0040 ms
12,393x faster
write every value
repetitive
30.9 ms
10.3 ms
3.0x faster
high-entropy
29.1 ms
12.4 ms
2.4x faster
The download and the first read, which is what a page that fetches one document waits for.
50,000 rows, fast 4G (9 Mbit/s, 85 ms)
data
JSON
jalapenojson
difference
receive the document
repetitive
659 ms
488 ms
1.4x faster
high-entropy
2,163 ms
1,805 ms
1.2x faster
receive it and read every value
repetitive
699 ms
525 ms
1.3x faster
high-entropy
2,215 ms
1,848 ms
1.2x faster
receive it and read one number column
repetitive
718 ms
491 ms
1.5x faster
high-entropy
2,245 ms
1,808 ms
1.2x faster
receive it and read one row
repetitive
702 ms
490 ms
1.4x faster
high-entropy
2,216 ms
1,807 ms
1.2x faster
Warmed up: json.loads and json.dumps against decode(), view() and encode().
50,000 rows
data
JSON
jalapenojson
difference
read every value
repetitive
76.2 ms
46.6 ms
1.6x faster
high-entropy
73.9 ms
69.3 ms
1.1x faster
read every value, as columns
repetitive
140 ms
19.7 ms
7.1x faster
high-entropy
135 ms
31.0 ms
4.3x faster
read one number column
repetitive
76.8 ms
0.679 ms
113x faster
high-entropy
71.6 ms
0.678 ms
106x faster
read one text column
repetitive
88.9 ms
6.61 ms
13x faster
high-entropy
81.7 ms
5.74 ms
14x faster
read one row
repetitive
81.1 ms
0.033 ms
2,423x faster
high-entropy
74.2 ms
0.034 ms
2,201x faster
write every value
repetitive
64.1 ms
45.4 ms
1.4x faster
high-entropy
69.5 ms
50.2 ms
1.4x faster
cJSON and yyjson against jj_parse() and its accessors.
50,000 rows
data
cJSON
yyjson
jalapenojson
vs cJSON
vs yyjson
read every value
repetitive
63.1 ms
9.79 ms
3.58 ms
18x faster
2.7x faster
high-entropy
64.9 ms
10.7 ms
5.06 ms
13x faster
2.1x faster
read one number column
repetitive
62.6 ms
9.46 ms
0.148 ms
422x faster
64x faster
high-entropy
64.8 ms
10.1 ms
0.148 ms
436x faster
68x faster
read one text column
repetitive
64.1 ms
9.27 ms
1.10 ms
58x faster
8.4x faster
high-entropy
62.5 ms
10.5 ms
1.19 ms
53x faster
8.8x faster
read one row
repetitive
61.2 ms
8.94 ms
0.00036 ms
168,563x faster
24,640x faster
high-entropy
60.6 ms
9.60 ms
0.00037 ms
163,319x faster
25,887x faster
Was not measured for this version, because the machine that ran the benchmark has clang 18.1.3 and wasm-ld, but not wasi-libc or compiler-rt's wasm32 builtins, which the two WebAssembly modules are built against, so neither could be built.
Measured on 2026-10-07 at commit 5548d6a: Intel Xeon Processor @ 2.80GHz, 4 cores, Linux 6.18.44-fc-v77; Node 22.22.0, Python 3.13.16, cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 with -O2, cJSON 1.7.19, yyjson 0.13.0. The full report also has results for 1,000 rows, the first call in a new process, brotli and a slow network, and describes how each number was measured.
import { encode, decode } from "./jalapenojson.js";
const buf = encode(rows, [["id","i",4], ["name","s",16], ["seen","z"]]);
const back = decode(buf); // Uint8Array in, objects out
back[0].seen // { instant: Date, offsetMinutes: 120, micros: 0 }
JavaScript
import { view } from "./jalapenojson.js";
const t = view(buf);
t.rows // how many rows
t.get(0, "name") // one value
t.column("name") // one column, as an array
t.numbers("price") // a number column, as a typed array
t.row(0) // a whole row
t.get(0, "items") // a nested list, as a view over its rows
C
jj_doc doc;
int rc = jj_parse(buf, len, &doc); // buf must outlive doc
if (rc != JJ_OK) return fprintf(stderr, "%s\n", jj_strerror(rc));
long long id = jj_int(jj_field(&doc, 0, 0));
jj_free(&doc);
The guide has more, including dates, booleans, nested lists, calculating column sizes and the rest of the C API.
Download
Each implementation is a single file with no dependencies. Copy it into your project.