B-FAST is an ultra-high performance binary serialization protocol, developed in Rust for Python and TypeScript ecosystems. It's designed to replace JSON in critical routes where latency, CPU usage, and bandwidth are bottlenecks.
"Performance is not just about speed—it's about efficiency where it matters most"
B-FAST was born from the recognition that modern applications need more than just fast serialization—they need smart serialization that adapts to real-world constraints. After extensive optimization, B-FAST operates in the sub-microsecond realm (676 ns encode / 754 ns decode for 100 objects), achieving 4.1x faster than orjson for objects, 5.7x faster on slow networks, and over 3,100,000 frames/s in streaming protocol.
Philosophy: We believe that the future of data transfer lies not in raw CPU speed alone, but in intelligent protocols that minimize network overhead while maintaining excellent performance. B-FAST represents our contribution to a more efficient, bandwidth-conscious web.
Full documentation available at: https://marcelomarkus.github.io/b-fast/
- Rust Engine: Native serialization without Python interpreter overhead.
- Sub-Microsecond Latency: Encodes 100 structured objects in 676 ns and decodes in 754 ns.
- Pydantic Native: Reads Pydantic model attributes directly from memory, skipping the slow .model_dump() process.
- Zero-Copy NumPy: Serializes tensors and numeric arrays directly, achieving 14-96x speedup vs JSON/orjson.
- Parallel Compression: LZ4 with multi-thread processing for large payloads (>1MB).
- Cache Optimized: Aligned allocation, direct C-API lists, interned strings, and zero intermediate copies.
| Operation | Time (ns) | Equivalent Ops / Second | Speedup vs Baseline |
|---|---|---|---|
| Encode (100 objects) | 676 ns | > 1,470,000 ops/s | 🚀 2.1x faster |
| Decode (100 objects) | 754 ns | > 1,320,000 ops/s | 🚀 2.6x faster |
| Format | Time (ms) | Speedup |
|---|---|---|
| JSON | 12.0ms | 1.0x |
| orjson | 8.19ms | 1.5x |
| B-FAST | 2.01ms | 🚀 6.0x |
B-FAST is 4.1x faster than orjson!
| Metric | Performance | Speedup / Throughput |
|---|---|---|
| Streaming Decode (Aligned) | 0.31ms (314µs) | ~3,180,000 frames/s (145x vs NDJSON) |
| Streaming Decode (Fragmented) | 0.32ms (322µs) | ~3,100,000 frames/s (Zero TCP penalty) |
| Single Frame Latency | 2.0ns | Real-time instant parsing |
| Sustained Stream Throughput | > 3,100,000 frames/s | Ultra-high-frequency telemetry & AI feeds |
Complete test including network transfer and deserialization (10,000 objects):
| Format | Total Time | Speedup vs orjson |
|---|---|---|
| JSON | 114.5ms | 0.8x |
| orjson | 91.7ms | 1.0x |
| B-FAST + LZ4 | 16.1ms | 🚀 5.7x |
| Format | Total Time | Speedup vs orjson |
|---|---|---|
| JSON | 29.4ms | 0.5x |
| orjson | 15.3ms | 1.0x |
| B-FAST + LZ4 | 7.2ms | 🚀 2.1x |
| Format | Total Time | Speedup vs orjson |
|---|---|---|
| JSON | 20.9ms | 0.4x |
| orjson | 7.7ms | 1.0x |
| B-FAST + LZ4 | 6.3ms | 🚀 1.2x |
- ⚡ Microservices & High-Frequency Trading: Sub-microsecond latency (< 800 ns per 100 objects)
- 🌊 Real-time Streaming & AI Telemetry: > 3,100,000 frames/s with zero TCP fragmentation penalty
- 📱 Mobile/IoT: 89% data savings + 5.7x performance on slow networks
- 🌐 APIs with slow networks: Up to 5.7x faster than orjson
- 📊 Data pipelines: 14-96x speedup for NumPy arrays
- 🗜️ Storage/Cache: Superior integrated compression
- 🚀 Simple objects: 4.1x faster than orjson
# Basic installation
pip install bfast-py
# With FastAPI support
pip install "bfast-py[fastapi]"or with uv:
uv add bfast-py
# or
uv add "bfast-py[fastapi]"npm install bfast-clientB-FAST supports two integration styles for FastAPI:
Zero-code route refactoring. Existing routes return JSON by default, but automatically return compressed B-FAST binary when requested by clients via Accept: application/x-bfast:
from fastapi import FastAPI
from b_fast.fastapi import BFastMiddleware
app = FastAPI()
app.add_middleware(BFastMiddleware, compress=True)
@app.get("/users")
def get_users():
# Returns standard JSON to browsers
# Automatically returns B-FAST binary when requested via Accept: application/x-bfast!
return [{"id": i, "name": f"User {i}"} for i in range(1000)]Explicit route control without middleware:
from fastapi import FastAPI
from pydantic import BaseModel
from b_fast import BFastResponse, BFastStreamingResponse
app = FastAPI()
class User(BaseModel):
id: int
name: str
# 1. Explicit B-FAST binary response
@app.get("/users", response_class=BFastResponse)
async def get_users():
# Returns binary B-FAST data with automatic LZ4 compression
return [User(id=i, name=f"User {i}") for i in range(1000)]
# 2. ⚡ Streamable HTTP (Progressive Chunks, 3.18M frames/sec)
@app.get("/users/stream")
async def stream_users():
async def user_generator():
for i in range(1000):
yield User(id=i, name=f"User {i}")
# Streams framed chunks with Content-Type: application/x-bfast-stream
return BFastStreamingResponse(user_generator())from ninja import NinjaAPI
from b_fast.django import BFastRenderer, BFastHttpResponse
# Django Ninja with BFastRenderer
api = NinjaAPI(renderer=BFastRenderer())
@api.get("/users")
def get_users(request):
return [{"id": i, "name": f"User {i}"} for i in range(1000)]
# Standard Django View
def django_view(request):
return BFastHttpResponse({"status": "ok"})from b_fast import BFast, encode_dataframe
import polars as pl
df = pl.DataFrame({"id": [1, 2, 3], "score": [95.0, 88.0, 92.5]})
# Direct native serialization in BFast
packed = BFast().encode_packed(df, compress=True)
# Or with orientation control ('records', 'columns', 'split')
col_data = encode_dataframe(df, orient="columns")from b_fast import FastMCPBFast, bfast_tool
mcp = FastMCPBFast("data-service")
@mcp.tool()
@bfast_tool()
def query_records(limit: int = 100):
return [{"id": i, "metric": i * 1.5} for i in range(limit)]import { bfastFetch, bfastQueryOptions } from 'bfast-client';
import { useQuery } from '@tanstack/react-query';
import { z } from 'zod';
const UserSchema = z.object({ id: z.number(), name: z.string() });
type User = z.infer<typeof UserSchema>;
// Direct Fetch
const users = await bfastFetch<User[]>('/users');
// In React with TanStack Query
function UserComponent() {
const { data: user } = useQuery(
bfastQueryOptions<User>({
queryKey: ['user', 1],
url: '/users/1',
schema: UserSchema, // Runtime schema validation
})
);
return <div>{user?.name}</div>;
}import { decodeReadableStream } from 'bfast-client';
async function streamData() {
const response = await fetch('/users/stream');
// Iterates over incoming network frames in real-time
for await (const user of decodeReadableStream(response.body!)) {
console.log('Received user in real-time:', user);
}
}B-FAST provides a standardized, curated llms.txt endpoint for AI tools (OpenCode, Cursor, Claude Code, ChatGPT, Windsurf, and GitHub Copilot).
Prompt your AI assistant directly:
Follow the B-FAST guidelines at https://marcelomarkus.github.io/b-fast/llms.txt to implement binary endpoints.Key Achievements:
- 🚀 4.1x faster than orjson for simple objects (2.01 ms)
- 🚀 5.7x faster than orjson on 100 Mbps networks (round-trip)
- 🌊 12,500+ frames/sec sustained streaming throughput (~139 µs latency)
- 📦 89% smaller payloads with built-in LZ4 compression
- ⚡ 14-96x speedup for NumPy arrays
- 🎯 Competitive even on ultra-fast 10 Gbps networks
B-FAST performance comparison across 6 key scenarios: simple objects encoding, zero-copy NumPy arrays, payload size, 100 Mbps round-trip, streaming decode time, and streaming throughput. B-FAST demonstrates clear superiority in speed (2.1-14x faster) and bandwidth efficiency (90% reduction with LZ4).
Developed by: marcelomarkus
Distributed under the MIT License. See LICENSE for more information.
