Skip to content

Latest commit

 

History

History
943 lines (687 loc) · 35.7 KB

File metadata and controls

943 lines (687 loc) · 35.7 KB

rumi

  • Specification 0.1.0
  • Status Draft
  • Date 2026-09-24
  • License GPLv3

rumi is stateless raster storage for AI4EO. Its GeoTIFF-inspired format stores compressed raster tensors with up to four dimensions. A file contains either an Image (B, Y, X) or a temporal Cube (T, B, Y, X).

The format has the following properties:

  • Stateless reads. The external binary header locates every compressed frame without parsing the .rumi file or retaining state between requests.
  • OpenZL compression. Stores each frame as an independent, self-contained OpenZL frame.
  • Predictable layout. Files with the same band and frame counts begin their frame data at the same byte.
  • Described bands and time. Names every band and labels every time step in a trailer after the frame data.
  • Canonical structure. Restricts the file to one fixed IFD, one frame order, and contiguous frame data.
  • Fixed-size georeferencing. Stores an EPSG CRS and affine transform in a 160-byte block derived from GeoTIFF tags, with a defined value for ungeoreferenced rasters.

The filename extension for the format is .rumi.

A raster array is encoded as a rumi file and a rumi header blob

The key words MUST, MUST NOT, SHOULD, and MAY are to be interpreted as described in RFC 2119 when, and only when, they appear in capitals.

Scope

This document defines:

  • the rumi data model and frame layouts;
  • the GeoTIFF-inspired rumi file profile;
  • the band descriptions and time coordinates every file carries; and
  • the binary layout of the external rumi header blob.

It does not define the OpenZL frame format or a catalogue format.

This document contains everything needed to implement a rumi reader. The only supported writer is the one provided by rumi; independent writers are outside the compatibility policy.

Unless a section says otherwise, all integer arithmetic used to validate or derive sizes, counts, and offsets is exact. A reader MUST reject an input when a required result cannot be represented by its implementation.

All multi-byte numeric values defined by rumi are little-endian. This includes the file header, directory entries and values, decoded sample components, the trailer, and the external header blob. OpenZL defines the bytes inside a compressed frame; after decoding, each multi-byte sample component is little-endian. An API may convert decoded samples to the host's native byte order.

Selective read model

This section is informative. A selective read combines band and time positions, a spatial window, the external header that belongs to the file, and access to the .rumi file.

  1. The window identifies the tile rows and columns it intersects.
  2. frame_unit determines which band and time positions are part of each frame index and which are decoded inside a frame.
  3. The header's packed byte counts reconstruct the offset and length of every required frame.
  4. Each selected range is decoded as an independent OpenZL frame and its requested samples are placed in the output array.

For example, consider a Cube with shape (T=3, B=4, Y=1024, X=1024), nominal tiles of 256 × 256, and frame_unit = 0. Its spatial grid is 4 × 4, so g = 16 and N = g * B * T = 192.

A request for time position t = 2, bands b = 1 and b = 3, and the window y = [512, 768), x = [256, 512) covers the tile at row = 2, col = 1.

spatial = row * tiles_across + col
        = 2 * 4 + 1
        = 9

index(b=1, t=2) = (9 * 4 + 1) * 3 + 2 = 113
index(b=3, t=2) = (9 * 4 + 3) * 3 + 2 = 119

A spatial, band, and time selection becomes two frame ranges

The reader reconstructs the offsets of frames 113 and 119 as defined in Offset reconstruction, fetches those two ranges, and leaves the other 190 frames untouched.

Data model

rumi uses the following data model.

Deep Learning Raster Data Model

term shape definition
tile (h, w) one band and time step at one tile location
cell (B, h, w) or (T, B, h, w) all samples at one tile location
Image (B, Y, X) one raster grid; a rumi file with T = 1
Cube (T, B, Y, X) T time-ordered, grid-aligned Images in one file
ImageCollection — a set of Images that need not share a grid
CubeCollection — a set of Cubes that need not share a grid

T, B, Y, and X denote time count, band count, image length, and image width. h and w denote the actual dimensions of a tile; they may be smaller than the nominal tile dimensions at the image boundary.

The bands and time steps of an Image or Cube are described in the Trailer.

Collections are represented outside the file, for example by a catalogue of rumi files. Their representation is out of scope.

Frames

tile and cell belong to the logical raster model. A frame is the physical unit of compression and random access. Each frame is a self-contained OpenZL frame that carries its own graph and codec parameters; neither is stored in the IFD or external header. Its decoded shape and sample order are specified by frame_unit.

A frame MUST contain exactly one of:

  • one tile for one (b, t) pair at one tile location; or
  • one cell at one tile location.

Tiles and cells are data model units; a frame is their compressed storage unit

frame_unit

frame_unit selects one of the decoded layouts below. b, t, h, and w mean band, time, height, and width. The rightmost axis changes fastest.

frame_unit decoded frame valid when
0 h w any B and T
1 b h w or t h w exactly one of B, T exceeds 1
2 h w b or h w t exactly one of B, T exceeds 1
3 b t h w B > 1 and T > 1
4 t b h w B > 1 and T > 1
5 b h w t B > 1 and T > 1
6 t h w b B > 1 and T > 1
7 h w b t B > 1 and T > 1
8 h w t b B > 1 and T > 1
9 h w B > 1 and T > 1

Units 1 and 2 place one non-spatial axis around h w. That axis is b when B > 1, and t when T > 1.

Units 0 and 9 differ only in the order the index walks the two axes at one tile location: b then t for 0, t then b for 9. That order is observable only when both B and T exceed 1, which is why 9 is valid nowhere else.

The diagram shows how the two frame types use the band and time axes. For a tile, b and t select the frame. For a cell, they are part of the decoded frame.

frame_unit keeps band and time outside a tile frame or places them around h w for cell storage

Within a decoded frame, h w MUST stay together and in that order.

The registry is complete and append-only; existing values MUST NOT be reassigned. A reader MUST reject any frame_unit value or (B, T) combination not listed above. A valid frame_unit determines the decoded sample order and the number of frames N; it does not otherwise change the file structure.

Choosing a frame unit

This section is informative. Units 0 and 9 provide the finest access: one frame contains one tile for one band and one time step. Every other unit places one cell in each frame, allowing OpenZL to model correlation between bands, time steps, or both.

A frame is the smallest unit a reader decodes, so the order of samples inside it changes only compression. A frame is a tile or a whole cell, never an intermediate group, and the index order of units 0 and 9 decides which tiles sit next to each other.

An axis before h w is stored as contiguous planes. An axis after h w is interleaved within each pixel. When both axes precede h w, the axis next to h w varies between adjacent planes.

The best unit depends on the expected reads and the data. Compression SHOULD be measured when more than one unit fits the access pattern.

Frame index

row and col select a tile location in the spatial grid. Tile locations are traversed in row-major order.

tiles_across = ceil(image_width / tile_width)
tiles_down   = ceil(image_length / tile_length)
g            = tiles_across * tiles_down

When a frame holds a cell, there is one frame per tile location.

frame_index(row, col) = row * tiles_across + col
N                     = g

When a frame holds a tile, the index also walks the band and time axes.

spatial = row * tiles_across + col

frame_unit 0:  frame_index = (spatial * B + b) * T + t
frame_unit 9:  frame_index = (spatial * T + t) * B + b

N = g * B * T

Validating frame_unit

Tag 65000 stores frame_unit in the file. Its value MUST be registered and valid for B and T.

The entry counts of TileOffsets and TileByteCounts MUST also match the frame unit:

count == g * B * T   when the frame holds a tile
count == g           when the frame holds a cell

A reader MUST reject a file when either count is wrong. The counts do not distinguish 0 from 9, or one full-frame layout from another; tag 65000 does.

Frame order

Frames MUST be stored in increasing frame_index order. TileOffsets and TileByteCounts MUST use the same order.

This order allows the external header to reconstruct offsets with a prefix sum. A reader MUST reject a file whose TileOffsets do not match the reconstructed offsets in frame-index order.

Sample encodings

sample_format gives the sample type and bits_per_sample its width in bits. Unsigned integers use ordinary binary representation, and signed integers use two's-complement representation. IEEE formats use the IEEE 754 binary16, binary32, or binary64 encoding named in the table. Boolean samples use one decoded byte whose value is 0 or 1, although their logical width is one bit.

A complex sample stores two equal-width components: real first, then imaginary. For complex formats, bits_per_sample is their combined width.

sample_format meaning
1 unsigned integer
2 signed integer
3 IEEE floating point
6 complex IEEE floating point
100..103, 107, 108 rumi-private ML floating point types

Only the following pairs are valid:

sample_format bits_per_sample decoded bytes encoding
1 1 1 boolean
1 8 1 unsigned 8-bit integer
1 16 2 unsigned 16-bit integer
1 32 4 unsigned 32-bit integer
1 64 8 unsigned 64-bit integer
2 8 1 signed 8-bit integer
2 16 2 signed 16-bit integer
2 32 4 signed 32-bit integer
2 64 8 signed 64-bit integer
3 16 2 IEEE 16-bit floating point
3 32 4 IEEE 32-bit floating point
3 64 8 IEEE 64-bit floating point
6 32 4 complex IEEE floating point, 16-bit components
6 64 8 complex IEEE floating point, 32-bit components
6 128 16 complex IEEE floating point, 64-bit components
100 8 1 float8 E4M3FN
101 8 1 float8 E5M2
102 16 2 bfloat16
103 8 1 float8 E8M0FNU
107 8 1 float8 E4M3FNUZ
108 8 1 float8 E5M2FNUZ

A reader MUST reject any pair not listed above.

The pairs (1, 2), (1, 4), (2, 2), (2, 4), (5, 32), (5, 64), (104, 6), (105, 6), and (106, 4) are reserved. The sample_format values 5, 104, 105, and 106 are also reserved and MUST NOT be assigned another meaning.

bfloat16 has one sign bit, eight exponent bits, and seven fraction bits, with the exponent and special values of IEEE binary32. E4M3FN, E4M3FNUZ, E5M2, and E5M2FNUZ use the ONNX float8 encodings. E8M0FNU uses the corresponding encoding in the OCP Microscaling Formats (MX) Specification 1.0.

All bands and time steps in a file MUST use the same pair.

bits_per_sample is the logical width of a sample, not necessarily its storage stride. Boolean samples occupy one decoded byte and every byte MUST be 0 or 1. Packed boolean storage MUST NOT be used.

The decoded frame size is defined by:

decoded_samples  = h * w           when the frame holds a tile
                   B * T * h * w   when the frame holds a cell

bytes_per_sample = the decoded bytes in the sample-encoding table

decoded_frame_bytes = decoded_samples * bytes_per_sample

An absent axis contributes a factor of one. Resource limits applies to decoded_frame_bytes.

DLPack representation

An implementation that exports decoded samples through DLPack MUST use the following (code, bits, lanes) values. Every tensor is native-endian CPU memory and one tensor element corresponds to one rumi sample. The float8 codes require DLPack 1.1 or newer.

encoding DLPack (code, bits, lanes)
signed integers (kDLInt=0, bits_per_sample, 1)
unsigned integers (kDLUInt=1, bits_per_sample, 1)
boolean (kDLBool=6, 8, 1)
IEEE floats (kDLFloat=2, bits_per_sample, 1)
complex IEEE floats (kDLComplex=5, bits_per_sample, 1)
bfloat16 (kDLBfloat=4, 16, 1)
float8 E4M3FN (kDLFloat8_e4m3fn=10, 8, 1)
float8 E4M3FNUZ (kDLFloat8_e4m3fnuz=11, 8, 1)
float8 E5M2 (kDLFloat8_e5m2=12, 8, 1)
float8 E5M2FNUZ (kDLFloat8_e5m2fnuz=13, 8, 1)
float8 E8M0FNU (kDLFloat8_e8m0fnu=14, 8, 1)

The file's logical width and DLPack's element width differ for boolean data; the decoded storage width connects them. A reader MUST NOT silently cast, reinterpret, or widen a sample to satisfy a consumer.

Bit-packed arrays

Time residuals and frame byte-count residuals use the same bit packing.

Values are stored consecutively using bits bits each. Bit j of value i occupies bit position i * bits + j, with bits numbered from the least significant bit of each byte. For n values, the region is exactly ceil(n * bits / 8) bytes. Unused bits in the final byte MUST be zero.

When bits is 0, the region is empty and every value is zero.

File profile

A file is rumi compliant when all of the following hold.

  • It begins with the rumi file header and contains exactly one rumi IFD.
  • It is tiled and has no overviews, masks, strips, or auxiliary IFDs.
  • Its IFD precedes the frame data, and its tags, values, and frame placement follow Fixed IFD, without gaps or padding.
  • Its sample encoding is listed in Sample encodings.
  • Its georeferencing follows Georeferencing.
  • Each frame is a self-contained OpenZL frame.
  • FrameUnit, TileOffsets, and TileByteCounts satisfy Validating frame_unit.
  • Every frame is present, every byte count is greater than zero, and the frames form one contiguous run in frame-index order.
  • It ends with the trailer defined in Trailer.

File header

Every rumi file begins with this 16-byte header.

offset size type name
0 4 bytes magic
4 2 uint16 version
6 2 uint16 reserved
8 8 uint64 ifd_offset

The magic bytes spell ASCII RUMI: 52 55 4D 49. The current version is 1, reserved is zero, and ifd_offset is 16. A reader MUST reject any other value.

The IFD uses the 20-byte entry layout and tag numbers derived from BigTIFF, but rumi defines its own tags, placement, and alignment. A rumi file is not a TIFF, BigTIFF, or GeoTIFF file.

Fixed IFD

A rumi IFD MUST contain exactly the tags below, in rising tag order. A reader that validates the file or builds an external header from it MUST reject any other tag.

B is samples_per_pixel, T is time_count, and N is the frame count.

The IFD begins with the eight-byte entry count 13 and ends with an eight-byte zero offset for the next IFD. Each entry has this layout:

offset size type name
0 2 uint16 tag
2 2 uint16 type
4 8 uint64 count
12 8 bytes value or offset

The type codes are 3 for SHORT, 4 for LONG, 12 for DOUBLE, and 16 for LONG8. These represent uint16, uint32, IEEE 754 binary64, and uint64, respectively.

tag name type count
256 ImageWidth LONG 1
257 ImageLength LONG 1
258 BitsPerSample SHORT B
277 SamplesPerPixel SHORT 1
322 TileWidth SHORT 1
323 TileLength SHORT 1
324 TileOffsets LONG8 N
325 TileByteCounts LONG N
339 SampleFormat SHORT B
34264 ModelTransformationTag DOUBLE 16
34735 GeoKeyDirectoryTag SHORT 16
65000 FrameUnit SHORT 1
65001 TimeCount LONG 1

Every file carries all 13 tags.

ImageWidth, ImageLength, TimeCount, TileWidth, TileLength, and SamplesPerPixel MUST be greater than zero.

BitsPerSample and SampleFormat MUST contain B repetitions of one pair listed in Sample encodings.

Placement

The IFD starts at byte 16, immediately after the rumi file header. Its size is 8 + 20 * 13 + 8 = 276 bytes: an eight-byte entry count, 13 entries, and an eight-byte zero offset for the next IFD.

Values of 8 bytes or less MUST be stored in the IFD entry. Larger values MUST follow the IFD in rising tag order, without gaps. Unused bytes in an inline value MUST be zero. For an external value, the last eight bytes of the entry store its uint64 file offset.

The frame data starts immediately after the last external value. Padding or alignment bytes MUST NOT be inserted in the external area.

Deriving base_frame_offset

The IFD size is fixed. Only values larger than 8 bytes contribute to the external area before the frames.

external = (2 * B  if B >= 5 else 0)      # 258 BitsPerSample
         + (8 * N  if N >= 2 else 0)      # 324 TileOffsets
         + (4 * N  if N >= 3 else 0)      # 325 TileByteCounts
         + (2 * B  if B >= 5 else 0)      # 339 SampleFormat
         + 128                            # 34264 ModelTransformationTag
         + 32                             # 34735 GeoKeyDirectoryTag

base_frame_offset = 16 + 276 + external
                  = 292 + external

The result MUST match the first entry of TileOffsets. Files with the same B and N start their frame data at the same byte.

Georeferencing

Every rumi file carries ModelTransformationTag and GeoKeyDirectoryTag. ModelPixelScaleTag (33550), ModelTiepointTag (33922), GeoDoubleParamsTag (34736), and GeoAsciiParamsTag (34737) MUST NOT appear. Files without georeferencing use the values in Undefined georeferencing.

ModelTransformationTag

ModelTransformationTag stores a 4 × 4 matrix as 16 doubles in row-major order. North-up rasters set the rotation terms to zero.

Given affine coefficients (x_res, row_rot, x_origin, col_rot, y_res, y_origin), the matrix is

x_res    row_rot  0  x_origin
col_rot  y_res    0  y_origin
0        0        0  0
0        0        0  1

The third row MUST be zero.

GeoKeyDirectoryTag

rumi represents a CRS by EPSG code. GeoKeyDirectoryTag contains a four-short header followed by three four-short keys, for a total of 16 shorts.

key id value
GTModelTypeGeoKey 1024 1 projected or 2 geographic
GTRasterTypeGeoKey 1025 1 PixelIsArea or 2 PixelIsPoint
GeographicTypeGeoKey or ProjectedCSTypeGeoKey 2048 or 3072 EPSG code

The four-short header MUST be (1, 1, 0, 3). Each key is stored as (key_id, 0, 1, value), in the order shown above.

A geographic CRS uses key 2048; a projected CRS uses key 3072. The key MUST agree with GTModelTypeGeoKey. The EPSG code MUST be defined by the EPSG registry and fall between 1024 and 32766.

A writer MUST obtain the CRS type from the EPSG registry or a PROJ database. It MUST NOT infer the type from the numeric code. A reader obtains the type from GTModelTypeGeoKey.

Undefined georeferencing

An ungeoreferenced file still carries both georeferencing tags.

ModelTransformationTag uses the following matrix.

1  0  0  0
0  1  0  0
0  0  0  0
0  0  0  1

GeoKeyDirectoryTag uses the same header and key representation:

key id value
GTModelTypeGeoKey 1024 0
GTRasterTypeGeoKey 1025 1 or 2
GeographicTypeGeoKey 2048 0

A reader MUST interpret GTModelTypeGeoKey = 0 as no CRS and MUST ignore the transformation matrix.

The two tags always occupy 160 bytes.

No other CRS representation is permitted. This excludes WKT, PROJ strings, ESRI codes, user-defined CRS values, engineering, compound and vertical CRSs, and coordinate epochs.

Trailer

Every rumi file ends with one trailer. It begins immediately after the last frame, and the file ends immediately after it. The trailer describes every band and labels every time step.

The trailer follows the frame data so that its size never moves a frame. A reader MUST NOT use it to locate, decode, or convert frames. The external header contains no band descriptions or time coordinates.

+----------------+---------------+-------------+-------------------------------+
| magic, version | band_texts[B] | time fields | time_residuals[C]             |
+----------------+---------------+-------------+-------------------------------+
  6 bytes          S bytes         22 bytes      ceil(C * time_bits / 8) bytes

S is the size of the band texts, defined in Band descriptions. The trailer size MUST be exactly 28 + S + ceil(C * time_bits / 8) bytes and MUST contain no padding. All multi-byte fields are little-endian.

Magic and version

offset size type name
0 4 uint32 magic
4 2 uint16 version

The four magic bytes spell ASCII TAIL: 54 41 49 4C. Read as a little-endian uint32, they equal 0x4C494154. A reader MUST reject any other value.

The current trailer version is 1. A reader that implements version 1 MUST reject any other value.

Band descriptions

The version is followed by one text for each of the B bands, in band order. Each text is stored as its length in bytes followed by the bytes themselves.

size type name
2 uint16 n[b]
n[b] bytes text
S = sum(2 + n[b])   for 0 <= b < B

Each text MUST be valid UTF-8, at least one byte long, and free of the byte 0x00. Two texts in one file MUST NOT be equal as byte sequences. A reader MUST reject a trailer whose texts break these rules.

This paragraph is informative. The recommended text gives the band name, a short description, and the wavelength, as in this text for the Sentinel-2 red band:

B4, Red, 664.5nm (S2A) / 665nm (S2B)

Time coordinates

The time fields follow the last band text. Every time step has a coordinate. An Image without a single acquisition instant, such as a DEM or an annual composite, uses an interval.

Let C be the number of coordinates stored in the trailer:

C = T       when time_type is 2 (instant)
C = 2 * T   when time_type is 1 (interval)

For instants, time(i) is the coordinate of time step i.

For intervals, time step i covers [time(2i), time(2i + 1)). Each step stores its own start and end, so intervals may leave gaps.

Offsets are relative to the first byte after the last band text.

offset size type name
0 1 uint8 time_type
1 1 uint8 time_bits
2 8 int64 time_epoch
10 8 int64 time_step
18 4 uint32 time_scale

time_type

value meaning
1 interval; each step is bounded by two coordinates
2 instant; the step is a point in time

A reader MUST reject any other value.

Instant coordinates MUST be non-decreasing.

Interval coordinates MUST satisfy:

time(2i) < time(2i + 1)         for 0 <= i < T
time(2i + 1) <= time(2i + 2)    for 0 <= i < T - 1

Intervals may meet or leave gaps, but MUST NOT overlap.

time_epoch, time_step and time_scale

time_scale is the number of seconds represented by one coordinate unit. It MUST be 86400 when every instant or interval endpoint is an exact whole-day offset from 1970-01-01T00:00:00Z; otherwise it MUST be 1. A reader MUST reject any other value.

Coordinate time(i) is a signed offset of time(i) * time_scale seconds from 1970-01-01T00:00:00Z; negative values represent times before that epoch. rumi follows POSIX time and does not represent leap seconds.

time_epoch MUST equal time(0). time_step is the slope of the prediction line used to encode the remaining coordinates:

time_step = 0                                      if C < 2
time_step = round((time(C-1) - time(0)) / (C-1))  otherwise

round chooses the nearest integer; exact halves round toward positive infinity.

time_bits

The number of bits used to encode each residual, as defined in Time residuals. It MUST be between 0 and 64.

Time residuals

After the time fields, the trailer stores the C coordinates as residuals against a straight line through time_epoch and time_step.

An axis whose coordinates lie on this line requires no packed region. Otherwise, the trailer stores their deviations from the line.

Encoding

Let time(i) be coordinate i in time_scale units.

predicted(i) = time_epoch + i * time_step
residual(i)  = time(i) - predicted(i)

A residual MUST fit in int64. It is mapped to an unsigned integer with zigzag encoding:

zigzag(x) = 2 * x       if x >= 0
            -2 * x - 1  otherwise

packed(i) = zigzag(residual(i))
time_bits = bit_length(max(packed))

bit_length(0) is 0.

A writer MUST use the minimum time_bits that represents the largest packed residual. residual(0) is always zero.

A reader MUST reject a trailer unless time_scale satisfies the rule above and the decoded coordinates reproduce the recorded time_epoch, time_step, and minimum time_bits. Every decoded coordinate MUST fit in int64 and satisfy the ordering required by time_type.

Packing

The C packed values are stored as defined in Bit-packed arrays, at time_bits bits each.

Decoding

residual(i) = packed(i) >> 1              if packed(i) is even
              -((packed(i) >> 1) + 1)     otherwise
time(i)     = time_epoch + i * time_step + residual(i)

Header blob

The header blob is a binary record stored outside the rumi file. It contains the raster fields and frame byte counts needed to locate a frame without parsing the file.

The blob consists of a fixed 32-byte header followed by packed frame byte counts. All multi-byte fields are little-endian.

+---------------+----------------------------------+
| Header        | frame_byte_counts[N]             |
+---------------+----------------------------------+
  32 bytes        ceil(N * count_bits / 8) bytes

N is derived as defined in Frame index.

The blob size MUST be exactly 32 + ceil(N * count_bits / 8) bytes and contain no padding.

The blob stores no offsets. They are reconstructed as defined in Deriving base_frame_offset and Offset reconstruction.

Header fields

offset size type name
0 4 uint32 magic
4 2 uint16 version
6 4 uint32 image_width
10 4 uint32 image_length
14 4 uint32 time_count
18 2 uint16 tile_width
20 2 uint16 tile_length
22 2 uint16 samples_per_pixel
24 1 uint8 bits_per_sample
25 1 uint8 sample_format
26 1 uint8 frame_unit
27 4 uint32 count_min
31 1 uint8 count_bits

magic

The magic value is 0x45564F4C, represented on the wire as 4C 4F 56 45. A reader MUST reject any other value.

version

The current binary format version is 1. A reader that implements version 1 MUST reject any other value.

image_width, image_length and time_count

The raster dimensions. Width and length are in pixels and match ImageWidth and ImageLength. time_count matches the value of TimeCount.

All three MUST be greater than zero. A time_count of 1 is an Image; anything larger is a Cube.

tile_width and tile_length

The nominal tile dimensions in pixels. They match TileWidth and TileLength.

Both values MUST be greater than zero.

The grid size is defined in Frame index.

Edge tiles are clipped to the image bounds and are not padded. A reader MUST derive their dimensions from the image shape and grid position.

samples_per_pixel

The band count B. This matches SamplesPerPixel in the IFD and MUST be at least 1.

bits_per_sample and sample_format

The sample type and width, as defined in Sample encodings.

frame_unit

The frame layout, as defined in frame_unit. Together with the image shape it determines N. Tag 65000 carries the same value inside the file.

count_min and count_bits

The frame byte count encoding, as defined in Frame byte counts. count_bits MUST be between 0 and 32.

Frame byte counts

After the fixed header, the blob stores the compressed size of each frame in frame-index order. Every count MUST be greater than zero. Each count is encoded as a residual from the minimum count.

Encoding

Let c[i] be the byte count of frame i.

count_min  = min(c)
count_bits = 0 if max(c) == count_min else bit_length(max(c) - count_min)

bit_length(x) is the minimum number of bits required to represent x.

A writer MUST use these values. A reader MUST reject a blob if count_min is not the minimum decoded count or count_bits is not the minimum required width.

A file with one frame always has count_bits = 0.

Packing

The N residuals c[i] - count_min are stored as defined in Bit-packed arrays, at count_bits bits each.

Decoding

c[i] = count_min + residual[i]

Every reconstructed count MUST fit in uint32.

Creating the header blob

A writer or header builder MUST create the blob from the finalized rumi file. Every duplicated value MUST match:

header blob rumi IFD
image_width ImageWidth
image_length ImageLength
time_count TimeCount
tile_width TileWidth
tile_length TileLength
samples_per_pixel SamplesPerPixel
bits_per_sample every BitsPerSample value
sample_format every SampleFormat value
frame_unit FrameUnit
frame_byte_counts[i] TileByteCounts[i]

frame_byte_counts and TileByteCounts MUST each contain N entries. The writer or header builder MUST NOT produce the blob if any comparison fails.

This check happens when the blob is created. A stateless reader can then treat the blob as authoritative and does not need to read or compare the IFD, TileOffsets, TileByteCounts, the trailer, or the total file size before reading a frame. How an application keeps a blob associated with its file is outside the scope of this specification.

Offset reconstruction

The blob stores no frame offsets. A reader derives base_frame_offset and reconstructs the offsets in frame-index order.

offset[0]     = base_frame_offset
offset[idx+1] = offset[idx] + frame_byte_counts[idx]

Frame offsets are prefix sums over the frame byte counts

Every reconstructed offset MUST fit in uint64.

A reader passes frame_byte_counts[idx] bytes at offset[idx] to the OpenZL decoder.

The byte just past the last frame is where the trailer begins.

trailer_offset = offset[N-1] + frame_byte_counts[N-1]

trailer_offset locates the band descriptions and time coordinates. Reading a frame does not require parsing the trailer.

The offset of frame k can also be expressed as

offset[k] = base_frame_offset
          + k * count_min
          + sum(residual[0:k])

Resource limits

A reader MUST check derived sizes before allocating memory or decoding a frame. This includes the band text lengths and the coordinate count C in the trailer. An operation that exceeds the reader's resource limits MUST fail before the allocation or decode. This does not make the rumi file invalid.

Changelog

  • 0.1.0. Initial draft.