Skip to main content

Binary Universal Language Kit 1.0
draft-thierry-bulk-08

Document Type Active Internet-Draft (individual)
Author Pierre Thierry
Last updated 2026-09-28
RFC stream (None)
Intended RFC status (None)
Formats
Stream Stream state (No stream defined)
Consensus boilerplate Unknown
RFC Editor Note (None)
IESG IESG state I-D Exists
Telechat date (None)
Responsible AD (None)
Send notices to (None)
draft-thierry-bulk-08
Network Working Group                                         P. Thierry
Internet-Draft                                               Comonad Dev
Intended status: Experimental                          27 September 2026
Expires: 31 March 2027

                   Binary Universal Language Kit 1.0
                         draft-thierry-bulk-08

Abstract

   This specification describes a simple, decentrally extensible and
   efficient format for data serialization.

Status of This Memo

   This Internet-Draft is submitted in full conformance with the
   provisions of BCP 78 and BCP 79.

   Internet-Drafts are working documents of the Internet Engineering
   Task Force (IETF).  Note that other groups may also distribute
   working documents as Internet-Drafts.  The list of current Internet-
   Drafts is at https://datatracker.ietf.org/drafts/current/.

   Internet-Drafts are draft documents valid for a maximum of six months
   and may be updated, replaced, or obsoleted by other documents at any
   time.  It is inappropriate to use Internet-Drafts as reference
   material or to cite them other than as "work in progress."

   This Internet-Draft will expire on 31 March 2027.

Copyright Notice

   Copyright (c) 2026 IETF Trust and the persons identified as the
   document authors.  All rights reserved.

   This document is subject to BCP 78 and the IETF Trust's Legal
   Provisions Relating to IETF Documents (https://trustee.ietf.org/
   license-info) in effect on the date of publication of this document.
   Please review these documents carefully, as they describe your rights
   and restrictions with respect to this document.  Code Components
   extracted from this document must include Revised BSD License text as
   described in Section 4.e of the Trust Legal Provisions and are
   provided without warranty as described in the Revised BSD License.

Table of Contents

   1.  Introduction  . . . . . . . . . . . . . . . . . . . . . . . .   4

Thierry                   Expires 31 March 2027                 [Page 1]
Internet-Draft                    BULK1                   September 2026

     1.1.  Rationale . . . . . . . . . . . . . . . . . . . . . . . .   4
       1.1.1.  Definitions . . . . . . . . . . . . . . . . . . . . .   4
       1.1.2.  State of the art  . . . . . . . . . . . . . . . . . .   6
       1.1.3.  Use cases . . . . . . . . . . . . . . . . . . . . . .   8
     1.2.  Format overview . . . . . . . . . . . . . . . . . . . . .  11
     1.3.  Conventions and Terminology . . . . . . . . . . . . . . .  12
   2.  BULK syntax . . . . . . . . . . . . . . . . . . . . . . . . .  13
     2.1.  Parsing algorithm . . . . . . . . . . . . . . . . . . . .  14
       2.1.1.  Summary of marker bytes . . . . . . . . . . . . . . .  15
       2.1.2.  Evaluation  . . . . . . . . . . . . . . . . . . . . .  15
     2.2.  Forms . . . . . . . . . . . . . . . . . . . . . . . . . .  17
       2.2.1.  Difference between sequence and form  . . . . . . . .  17
     2.3.  Atoms . . . . . . . . . . . . . . . . . . . . . . . . . .  17
       2.3.1.  nil . . . . . . . . . . . . . . . . . . . . . . . . .  17
       2.3.2.  Arrays  . . . . . . . . . . . . . . . . . . . . . . .  18
       2.3.3.  Reserved marker bytes . . . . . . . . . . . . . . . .  20
       2.3.4.  References  . . . . . . . . . . . . . . . . . . . . .  20
   3.  Kinds of namespaces . . . . . . . . . . . . . . . . . . . . .  21
   4.  BULK core namespace . . . . . . . . . . . . . . . . . . . . .  21
     4.1.  Version . . . . . . . . . . . . . . . . . . . . . . . . .  23
     4.2.  Namespaces and packages . . . . . . . . . . . . . . . . .  24
       4.2.1.  Importing a namespace . . . . . . . . . . . . . . . .  24
       4.2.2.  Importing a package . . . . . . . . . . . . . . . . .  25
       4.2.3.  Canonical identifiers . . . . . . . . . . . . . . . .  25
       4.2.4.  Namespace definition  . . . . . . . . . . . . . . . .  26
       4.2.5.  Package definition  . . . . . . . . . . . . . . . . .  28
       4.2.6.  Name definitions  . . . . . . . . . . . . . . . . . .  28
       4.2.7.  Mnemonic  . . . . . . . . . . . . . . . . . . . . . .  29
       4.2.8.  Explain . . . . . . . . . . . . . . . . . . . . . . .  29
     4.3.  Strings and other typed byte arrays . . . . . . . . . . .  29
       4.3.1.  Current encoding  . . . . . . . . . . . . . . . . . .  29
       4.3.2.  String  . . . . . . . . . . . . . . . . . . . . . . .  30
       4.3.3.  String with explicit encoding . . . . . . . . . . . .  30
       4.3.4.  IANA registered character set . . . . . . . . . . . .  30
       4.3.5.  Nested BULK stream  . . . . . . . . . . . . . . . . .  30
       4.3.6.  Blob  . . . . . . . . . . . . . . . . . . . . . . . .  31
     4.4.  Array operations  . . . . . . . . . . . . . . . . . . . .  31
       4.4.1.  Array concatenation . . . . . . . . . . . . . . . . .  31
       4.4.2.  Indexed data  . . . . . . . . . . . . . . . . . . . .  31
     4.5.  Booleans  . . . . . . . . . . . . . . . . . . . . . . . .  32
     4.6.  Substituton . . . . . . . . . . . . . . . . . . . . . . .  33
       4.6.1.  Substitution function . . . . . . . . . . . . . . . .  33
       4.6.2.  Examples  . . . . . . . . . . . . . . . . . . . . . .  33
       4.6.3.  Isolated scope  . . . . . . . . . . . . . . . . . . .  33
     4.7.  Arithmetic  . . . . . . . . . . . . . . . . . . . . . . .  34
       4.7.1.  Unsigned integer  . . . . . . . . . . . . . . . . . .  34
       4.7.2.  Signed integer  . . . . . . . . . . . . . . . . . . .  34
       4.7.3.  Fraction  . . . . . . . . . . . . . . . . . . . . . .  35

Thierry                   Expires 31 March 2027                 [Page 2]
Internet-Draft                    BULK1                   September 2026

       4.7.4.  Binary floating-point number  . . . . . . . . . . . .  35
       4.7.5.  Decimal floating-point number . . . . . . . . . . . .  35
     4.8.  Bytecodes . . . . . . . . . . . . . . . . . . . . . . . .  36
       4.8.1.  Prefix bytecode . . . . . . . . . . . . . . . . . . .  37
       4.8.2.  Postfix bytecode  . . . . . . . . . . . . . . . . . .  38
       4.8.3.  Arity definition  . . . . . . . . . . . . . . . . . .  39
   5.  Optimizing compactness  . . . . . . . . . . . . . . . . . . .  40
     5.1.  Packing . . . . . . . . . . . . . . . . . . . . . . . . .  40
     5.2.  Mixing literals . . . . . . . . . . . . . . . . . . . . .  41
     5.3.  Trade-offs  . . . . . . . . . . . . . . . . . . . . . . .  41
   6.  Profiles  . . . . . . . . . . . . . . . . . . . . . . . . . .  43
     6.1.  Profile redundancy  . . . . . . . . . . . . . . . . . . .  43
     6.2.  Standard profile  . . . . . . . . . . . . . . . . . . . .  43
     6.3.  Fixed BULK: all profile, no evaluation  . . . . . . . . .  44
   7.  Discovery of namespaces and packages  . . . . . . . . . . . .  44
   8.  Streaming . . . . . . . . . . . . . . . . . . . . . . . . . .  45
     8.1.  Broadcasting BULK . . . . . . . . . . . . . . . . . . . .  46
       8.1.1.  Server clipping . . . . . . . . . . . . . . . . . . .  46
       8.1.2.  Beacon expressions  . . . . . . . . . . . . . . . . .  46
       8.1.3.  Ogg encapsulation . . . . . . . . . . . . . . . . . .  47
       8.1.4.  Magrat encapsulation  . . . . . . . . . . . . . . . .  47
     8.2.  Full-duplex communication . . . . . . . . . . . . . . . .  48
       8.2.1.  Separate channels . . . . . . . . . . . . . . . . . .  48
       8.2.2.  Shared channel  . . . . . . . . . . . . . . . . . . .  48
   9.  Security Considerations . . . . . . . . . . . . . . . . . . .  50
     9.1.  Parsing . . . . . . . . . . . . . . . . . . . . . . . . .  50
     9.2.  Forwarding  . . . . . . . . . . . . . . . . . . . . . . .  50
     9.3.  Definitions . . . . . . . . . . . . . . . . . . . . . . .  51
     9.4.  Selectively parseable content . . . . . . . . . . . . . .  51
     9.5.  BULK formats and protocols  . . . . . . . . . . . . . . .  51
   10. IANA Considerations . . . . . . . . . . . . . . . . . . . . .  51
     10.1.  Media type . . . . . . . . . . . . . . . . . . . . . . .  52
       10.1.1.  application/bulk . . . . . . . . . . . . . . . . . .  52
       10.1.2.  text/bulk  . . . . . . . . . . . . . . . . . . . . .  53
     10.2.  Link relation  . . . . . . . . . . . . . . . . . . . . .  54
     10.3.  Ogg media mapping  . . . . . . . . . . . . . . . . . . .  54
   11. Acknowledgements  . . . . . . . . . . . . . . . . . . . . . .  54
   12. References  . . . . . . . . . . . . . . . . . . . . . . . . .  55
     12.1.  Normative References . . . . . . . . . . . . . . . . . .  55
     12.2.  Informative references . . . . . . . . . . . . . . . . .  56
   Appendix A.  Using the text notation as a format  . . . . . . . .  57
   Appendix B.  Robust namespace definition  . . . . . . . . . . . .  58
     B.1.  Complete authority  . . . . . . . . . . . . . . . . . . .  59
     B.2.  Selective authority . . . . . . . . . . . . . . . . . . .  59
     B.3.  Open authority  . . . . . . . . . . . . . . . . . . . . .  59
   Appendix C.  Forward compatibility  . . . . . . . . . . . . . . .  60
   Appendix D.  Arity-carrying forms . . . . . . . . . . . . . . . .  61
   Appendix E.  The difference between BULK and BULK formats . . . .  62

Thierry                   Expires 31 March 2027                 [Page 3]
Internet-Draft                    BULK1                   September 2026

     E.1.  Purposefully open: for generality . . . . . . . . . . . .  63
     E.2.  Purposefully limited: for safety  . . . . . . . . . . . .  63
   Appendix F.  Marking and extending media types with BULK  . . . .  64
   Author's Address  . . . . . . . . . . . . . . . . . . . . . . . .  65

1.  Introduction

1.1.  Rationale

   This specification aims at finding an original trade-off between
   transparency, syntax complexity, generality, extensibility,
   decentralization, discoverability, compactness, streamability,
   safety, processing speed and processing footprint for a data format
   (see definitions).  It is our opinion that every widely used existing
   format occupy a different position than this one in the solution
   space for formats, that none is better on all axes, and that this one
   is the current best on several axes, hence this new design.  It is
   also our opinion that some of those existing formats constitute an
   optimal solution for their specific use case, either in a absolute
   sense, or at least at the time of their design.  But the ever-
   changing field of IT now faces new challenges that call for a new
   approach.

   In particular, whereas the previous trend for Internet and Web
   standards and programming tools has been to create human-readable
   syntaxes for data and protocols, the advent of technologies like
   protocol buffers [protobuf], CBOR [RFC8949], Thrift [Thrift], the
   various binary serializations for JSON like Avro [Avro] or Smile
   [Smile], or the binary HTTP/2 [RFC7540] seem to indicate that the
   time is ripe for a generalized use of binary, reserved until now for
   the low-level protocols.  The lessons about flexibility learnt in the
   previous switch from binary to plain text can now be applied to
   efficient binary syntaxes.

1.1.1.  Definitions

   By transparency, we mean the property of a format that can be parsed
   even by an application that doesn't understand the semantics of every
   part of the processed data.

   By syntax complexity, we mean the number of different syntactic
   structures of the format.

Thierry                   Expires 31 March 2027                 [Page 4]
Internet-Draft                    BULK1                   September 2026

   High transparency and low syntax complexity mostly have value in the
   face of extension, as a fixed format doesn't need either, but even in
   that case, they might make the overall format and its implementation
   simpler, which can have a lot of value by reducing the risk of
   incompatible implementations and the likelihood that complex or
   opaque elements open up security vulnerabilities.

   Almost all extensible formats have a relatively high transparency and
   low syntax complexity for their extensible part.  The goal was thus
   to achieve a low syntax complexity for the whole format, while still
   having high generality (i.e. that extending the format without a new
   syntactic structure is as easy and practical as possible).

   A good counter-example is found in most programming languages.
   Adding a new branching construct cannot be done in a terse way
   without modifying the underlying implementation.  Such a construct
   either cannot be defined by user code (because of evaluation rules)
   or can in a terribly verbose and inconvenient way (with lots of
   boilerplate code).  Notable exceptions to this limitation of
   programming languages are Lisp languages (e.g. Common Lisp, Scheme or
   Clojure), languages with lazy evaluation (e.g. Haskell, Purescript or
   Agda) and stack (or concatenative) languages (e.g. Forth, Postscript
   or Factor).

   On the other hand, stack languages are the canonical examples of non-
   transparent formats.  Each operator takes a number of operands from
   the stack.  Not knowing the arity of an operator makes it impossible
   to continue parsing, even when its evaluation was optional to the
   final processing.  In the design space, stack languages completely
   sacrifice transparency to achieve one of the highest combination of
   extensibility, compactness and speed of processing.

   By generality, we mean the ability of a format to describe any type
   of data with a reasonable (or better yet, high) level of compactness
   and simplicity.  By analogy with data structures, while both arrays
   and linked lists are both able to store any kind of data, they
   actually do at the cost of complexity and transparency for arrays
   (they need the embedding of data structure in the data or in the
   processing logic) and size for linked lists (in-memory linked lists
   can waste as much as half or two third of the space for the overhead
   of the data structure).

   By extensibility, we mean the ability of a format to encode types and
   values that were not anticipated when the syntax was designed.

   By decentralization, we mean the ability to encode new types and
   values while avoiding name collisions, but without the need of
   coordination.  Note that the DNS, as we use it (e.g. in domain names

Thierry                   Expires 31 March 2027                 [Page 5]
Internet-Draft                    BULK1                   September 2026

   in module names in some programming languages, or in URIs in XML
   Namespaces), is _not_ decentralized in this sense, but distributed,
   as it cannot work without its root servers and prior knowledge of
   their location.

   By discoverability, we mean the ability for a processing application
   to automatically discover new extensions when it encounters their use
   in data, with no prior knowledge of them beforehand.

   By compactness, we mean the ability of a format to encode as many
   diverse types of data as possible with an overall size as small as
   possible.

   By streamability, we mean two levels.  The first, being streamed, is
   the ability of a format to be transported in pieces that can be
   processed before the next piece is available.  The second, being
   broadcasted, is, when a stream of data is cut in two at an arbitrary
   place, the ability of the second halt to be processed successfully.

   By safety, we mean the ability for a format to have a parser
   implementing the full specification while having a default behaviour
   that doesn't expose the system where it runs to attacks triggered by
   malicious input.

   By processing speed, we mean the property of a format that lends
   itself to benefit from current computing architectures to be
   processed at high speed overall (e.g. processing can be fast in
   itself, some shortcuts can be taken, or parallelization is possible).

   By processing footprint, we mean the ability of a format to be
   processed while using a low, possibly constrained quantity of memory.

1.1.2.  State of the art

   Transparency, generality and extensibility are usually highly-valued
   traits in formats design.  Programming languages obviously feature
   them foremost, although their generality usually stops at what they
   are supposed to express: procedures.  Most of them are ill-suited to
   represent arbitrary data, but notable exceptions include Lisp (where
   "code is data") and Javascript, from which a subset has been
   extracted to exchange data, JSON, which has seen a tremendous success
   for this purpose.  JSON may have some caveats with regards to
   generality and a relatively low compactness, but its design makes its
   parsing really straightforward and fast.  All of them, though, lack
   decentralization and discoverability.  Some of them make it possible
   to extend them in a distributed way if some discipline is followed
   (for example, by naming modules after domain names), but the
   discipline is not mandatory (and even with domain names, a change of

Thierry                   Expires 31 March 2027                 [Page 6]
Internet-Draft                    BULK1                   September 2026

   ownership makes it possible for name collisions).

   The SGML/XML family of formats also feature good transparency, syntax
   complexity, generality and extensibility and actually fare much
   better than programming languages on those axes.  XML namespaces also
   make XML naming distributed and there have been attempts at making it
   compact (e.g. EXI from W3C, Fast Infoset from ISO/ITU or EBML).

   All the previously cited formats clearly lack compactness, although
   just applying standard compression techniques would sacrifice only
   very little processing time to gain huge size reductions on most of
   their intended use cases, but compression may not address their
   ineffectiveness at storing arbitrary bytes (and compression of the
   base64 encoding of arbitrary bytes can be less efficient than
   compression of the arbitrary bytes).

   Neither JSON nor XML are suitable for streaming and the whole
   document must usually be parsed entirely before its content can be
   processed, which impacts processing speed and footprint.

   So-called binary formats pretty much exhibit the opposite trade-offs.
   Most of them have high syntax complexity, low generality and low
   extensibility to achieve better compactness.  Some are specifically
   designed for a great generality, but many lack extensibility.  When
   they are extensible, it's never in a decentralized way nor are they
   discoverable, both for reasons that have to do with compactness.
   They are usually extremely fast to parse, and while some are designed
   to be streamed, few can be broadcasted.

   Actually, many binary formats are not so much formats as they are
   formats frameworks, and exclude extensibility by design.  For each
   use case, an IDL compiler creates a brand new format that is
   essentially incompatible with all other formats created by the same
   compiler (EBML specifically cites this property among its own
   disadvantages).  If the IDL compiler and framework are well designed,
   such a format can represent an optimum in compactness and speed of
   processing, as the compiler can also automatically generate an ad-hoc
   optimized parser.

   Where extensibility has been planned in existing binary formats, it
   often doesn't get used that much or at all because of the
   complications around it.  Many binary formats include reserved values
   meant to extend them to future uses, like the CM field in the ZIP
   format.  A case like this one faces an chicken-and-egg problem: if
   you don't write and get a specification officially adopted,
   implementations might not want to include your extension, but if your
   extension is purely theoretical and hasn't been tested in the wild,
   you may face resistance to get it officially adopted.  This is

Thierry                   Expires 31 March 2027                 [Page 7]
Internet-Draft                    BULK1                   September 2026

   probably why even though most compression or compressed archive
   formats include the ability to later encode other compression
   methods, each new compression method usually comes with its own new
   format.

   When extensions are managed with any form of registry, another issue
   is that you usually need to reserve a large set of values for free
   experimentation, and once an extension gains any traction while in
   experimentation, its authors face the difficulty to switch all
   existing implementations to the definitive values they'll get.  And
   how experimenters choose their temporary values makes them vulnerable
   to conflicts with others.  Furthermore, the process of switching
   between the experimental and registered versions of the format or
   protocol might be error-prone and add a significant editorial
   workload ([I-D.bormann-cbor-draft-numbers] details some of the
   pitfalls and suggests a process to deal with this transition).

1.1.3.  Use cases

   Here are some cases where the use of BULK formats and protocols would
   make software engineers' and users' lives easier.

   +====================================+==============================+
   | Without BULK                       | With BULK                    |
   +====================================+==============================+
   | If a user has a huge collection    | The user's BULK image        |
   | of pictures but many of them       | software likely is able to   |
   | contain extended metadata that     | show the existence of        |
   | her image software doesn't         | different metadata in every  |
   | support yet, there is no safe and  | image file and, with safely  |
   | easy way for her to access it.     | auto-discovered data, can    |
   |                                    | visualize the extended       |
   |                                    | metadata, at least in a raw  |
   |                                    | form but with human-readable |
   |                                    | labels or, better yet,       |
   |                                    | mapped into a form it knows. |
   +------------------------------------+------------------------------+
   | If a user has a collection of      | Generic BULK tools can let   |
   | pictures with some in a file       | the user make queries about  |
   | format her image software doesn't  | her whole image collection,  |
   | support yet, she cannot access     | and even her broader file    |
   | the kind of common metadata that   | collection, making use of    |
   | she can expect most images to      | either common metadata       |
   | contain (like author or copyright  | formats between different    |
   | information).                      | image or file formats, or    |
   |                                    | the mapping between          |
   |                                    | different metadata formats   |
   |                                    | into a single one.  Queries  |

Thierry                   Expires 31 March 2027                 [Page 8]
Internet-Draft                    BULK1                   September 2026

   |                                    | could be for all objects,    |
   |                                    | whole files or entries       |
   |                                    | inside files, that have been |
   |                                    | flagged "confidential" or    |
   |                                    | belong to some entity, or    |
   |                                    | all images or videos where   |
   |                                    | some person has been tagged  |
   |                                    | as visible, for example.     |
   +------------------------------------+------------------------------+
   | If a new image or video            | If a new image or video      |
   | compression has been designed,     | compression has been         |
   | its designers usually need to      | designed, its designers      |
   | create a whole custom container    | don't need to create         |
   | format, with custom metadata       | anything more to embed it in |
   | format.  A lot of work is needed   | BULK.  If it has unusual     |
   | for the many image or video        | features, it is still pretty |
   | software to be able to             | easy to make what is         |
   | accommodate reading or writing     | backward-compatible with     |
   | this new container, new metadata   | existing data models fit     |
   | and new codec.  Because this new   | into the existing container  |
   | work involves parsing binary       | structure.  Even new kinds   |
   | data, it often is a source of      | of structures leverage       |
   | security vulnerabilities.          | existing parsing code,       |
   |                                    | minimizing attack surface.   |
   +------------------------------------+------------------------------+
   | If a new compression, signing or   | If a new compression,        |
   | encryption algorithm has been      | signing or encryption        |
   | designed, and its designers hope   | algorithm has been designed, |
   | to see it used in existing         | and its designers hope to    |
   | formats, they have to work         | see it used in existing BULK |
   | separately for each format where   | formats, their BULK          |
   | it may be used and then for each   | vocabulary is guaranteed to  |
   | software project implementing      | uniquely describe the data   |
   | them and in each case, they will   | produced by their algorithm  |
   | face a chicken-and-egg problem     | and can be readily embedded  |
   | where software projects may be     | in any existing BULK format. |
   | reluctant to put effort into       | The only work needed may be  |
   | something users may not            | to implement their algorithm |
   | ultimately want or benefit from.   | as a pure function that can  |
   | The designers may also need to     | be plugged into the BULK     |
   | interact with one or several       | evaluation mechanism, for    |
   | registries to get their algorithm  | the relevant programming     |
   | registered, which may be a         | languages, so that software  |
   | necessary upfront work.  Users     | projects just have to        |
   | need to wait for each software to  | register the link between    |
   | be updated to include the new      | BULK vocabulary and software |
   | algorithm and sometimes suffer     | plugin (or may not even need |
   | from incompatible implementations  | to intervene, if a safe      |

Thierry                   Expires 31 March 2027                 [Page 9]
Internet-Draft                    BULK1                   September 2026

   | at the format level (e.g. early    | mechanism is implemented to  |
   | adopters producing files with      | auto-discover software       |
   | obsolete numbers allocated to the  | plugins providing pure       |
   | algorithm).                        | functions).                  |
   +------------------------------------+------------------------------+
   | If a user encounters a file of an  | If a user encounters a BULK  |
   | unknown type, their system might   | file of a yet unknown type,  |
   | give them useful media type        | generic BULK tools can       |
   | identification or not, and based   | safely auto-discover human-  |
   | on that or file extension, an      | readable information about   |
   | advanced user might search in her  | the structure of the file,   |
   | system's software registry or on   | and parts of the file that   |
   | the Internet and find some         | the user can readily use can |
   | software that can visualize some   | be easily found and viewed   |
   | or all of the file's content.      | or extracted.                |
   +------------------------------------+------------------------------+
   | When a communication protocol is   | When a BULK-based            |
   | designed, its designers have to    | communication protocol is    |
   | choose a trade-off between         | designed, its designers can  |
   | latency caused by message size,    | make it both very compact    |
   | ease of debugging and complexity   | and easy to debug with very  |
   | of the parser.  It's deceptively   | little effort.  Generic      |
   | easy to make mistakes in the       | tools can present data on    |
   | syntax that make it harder to use  | the wire with the support of |
   | or implement, but using an         | safely auto-discovered data  |
   | existing syntax ties the protocol  | about the protocol.          |
   | with this syntax' limitations.     |                              |
   +------------------------------------+------------------------------+

                                  Table 1

   In a few of those cases, without BULK, a registry of executable
   software patches or plugins that can be automatically discovered
   could have been a solution, with the caveat that almost all of our
   current software architectures make that possibility more dangerous
   than it's worth (because of the broad ambient authority given to most
   code).

   With BULK, there is immediately a strong incentive to provide
   discoverability, at first for just data, not executable code, with
   safeties already in place on that process.  The data model of BULK
   then makes it more likely that code plugins are provided that are
   designed to work with extremely limited privileges, to act as
   transformers of BULK expressions.

Thierry                   Expires 31 March 2027                [Page 10]
Internet-Draft                    BULK1                   September 2026

1.2.  Format overview

   A BULK stream is a stream of 8-bit bytes, in big-endian order.
   Parsing a BULK stream yields a sequence of expressions, which can be
   either atoms or forms, which are sequences of expressions.

   Forms have a simple syntax: a starting byte marker, a sequence of
   expressions and an ending byte marker.

   Atoms each have a special syntax, for compactness purposes: they
   start with a marker byte, followed by a static or dynamic number of
   bytes, depending on the type.  But there are only 5 kinds: the nil
   atom, generic arrays, small arrays, small unsigned integers and
   references.

   Even booleans and floating-point numbers use this existing syntax
   without the need for special cases.

   References consist of a namespace marker (in almost all cases, a
   single byte) followed by an identifier within this namespace (a
   single byte).  All in all, a very little sacrifice is made in
   compactness for the benefit of a very simple and resilient syntax:
   apart from nil and small integers, nothing is smaller than 2 bytes,
   and as most forms involve a reference followed by some content, a
   form is usually 4 bytes + its content.

   A namespace marker in a BULK stream is associated to a namespace
   identified by some identifier guaranteed to be unique without
   coordination (like a UUID or cryptographical hash), thus ensuring
   decentralized extensibility.  The stream can be processed even if the
   application doesn't recognize the namespace.  Parsing remains
   possible thanks to the transparent syntax.

   Combination of BULK namespaces, BULK streams and even other formats
   doesn't need any content transformation to work.  Here are some
   examples:

   *  The content of a BULK stream, enclosed in form starting and ending
      byte markers, constitute a valid BULK expression.  Thus BULK
      streams can be aggregated or annotated within a BULK stream
      without modification.

   *  A BULK format could specify in its syntax the place for a metadata
      expression.  Whether the specification provides its own metadata
      forms or not, an application could use a BULK serialization for
      MARC, TEI Header, XML or RDF for this metadata expression.  The
      vocabulary selected would be univocally expressed by the namespace
      and every vocabulary would be parsed by the same mechanisms.

Thierry                   Expires 31 March 2027                [Page 11]
Internet-Draft                    BULK1                   September 2026

   *  Whenever a content must be stored as-is instead of serialized, or
      a highly-optimized ad hoc serialization exists for some data,
      anything can always be stored within an array.  They can contain
      arbitrary bytes and there is no limit to their size.

   Furthermore, BULK expressions can be evaluated.  Many expressions
   evaluate to themselves, but others evaluate to the result of
   executing a pure function, making it possible to serialize data in an
   even more compact form, by eliminating boilerplate data and repeated
   patterns.

1.3.  Conventions and Terminology

   The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
   "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this
   document are to be interpreted as described in [BCP14].

   Literal numerical values are provided in decimal or hexadecimal as
   appropriate.  Hexadecimal literals are prefixed with 0x to
   distinguish them from decimal literals.

   The text notation of the BULK stream uses mnemonics for some bytes
   sequences.  Mnemonics are sequences of characters, excluding all
   capital letters and white space, like this-is-one-mnemonic or what-
   the-%§!?#-is-that?. They are always separated by white space.
   Outside the use of mnemonics, a sequence of bytes (of one or more
   bytes) can be represented by its hexadecimal value as an unsigned
   integer prefixed by 0x (e.g. 0x3F or 0x3A0B770F).  Such a sequence of
   bytes can include dashes to make it more readable (e.g.
   0xDDA37D36-85E6-4E6D-9B51-959E1CCE366C).  Some types in this
   specification define a special syntax for their representation in the
   text notation.

   In the grammar, a shape is a pattern of bytes, following the rules of
   the text notation for a BULK stream.  Apart from mnemonics and fixed
   sequences of bytes, a shape can contain:

   *  an arbitrary sequence of a fixed number of bytes, represented by
      its size, i.e. a number of bytes in decimal immediately followed
      by a B uppercase letter (e.g. 4B)

   *  a typed sequence of bytes, represented by the name of its type, a
      capitalized word (e.g.  Foo); this means a sequence of bytes whose
      specific yield (see Parsing algorithm) has this type

Thierry                   Expires 31 March 2027                [Page 12]
Internet-Draft                    BULK1                   September 2026

   *  a named sequence of bytes (of zero or more bytes), represented by
      a sequence of any character excluding '{}' between '{' and '}'
      (e.g. {quux}); a named sequence can be typed or sized, in which
      case it is immediately followed by ':' and a type or size (e.g.
      {quux}:Bar or {quux}:12B)

   The shape that describes the byte sequence of an atom is called its
   parsing shape.  When a shape is given for a form, it merely describes
   the semantics of evaluating forms of that shape.  A reference used in
   such a shape can be used in different shapes, with unrelated
   semantics.

   For example, this specification defines a way do encode a string with
   explicit encoding with forms of the shape ( string {enc}:Expr
   {string}:Expr ).  But the shapes ( string {arg1}:Int {arg2}:Int ) or
   ( {arg1}:Int string {arg2}:Int ) are syntactically valid.  They just
   evaluate to themselves as lists of three expressions, as far as this
   specification is concerned.

2.  BULK syntax

   A BULK stream is a sequence of 8-bit bytes.  Bits and bytes are in
   big-endian order.  The result of parsing a BULK stream is a list of
   abstract data, called the abstract yield.  BULK parsing is injective:
   a BULK stream has only one abstract yield, but different BULK streams
   can have the same abstract yield (if they associate namespaces to
   different markers, see Namespaces and packages).

   A processing application is not expected to actually produce the
   abstract yield, but an adaptation of the abstract yield to its own
   implementation, called the concrete yield.  Also, some expressions in
   a BULK stream may have the semantics of a transformation of the
   abstract yield.  A processing application MAY thus not produce or
   retain the concrete yield but the result of its transformation.  This
   specification deals mainly with the byte sequence and the abstract
   yield and occasionally provides guidelines about the concrete yield.
   Of course, a processing application MAY not produce any concrete
   yield at all but produce various data structures and side effects
   from parsing the BULK stream (for example, an event sourced
   application may read its event log from a BULK stream and build its
   application state by applying the events, discarding each of them as
   soon as it has been applied).

   The abstract yield is a list of expressions.  Expressions can be
   atoms or forms.  Forms are lists of expressions.  If a byte sequence
   is parsed as an expression, this byte sequence is said to encode this
   expression.

Thierry                   Expires 31 March 2027                [Page 13]
Internet-Draft                    BULK1                   September 2026

   When a sequence of bytes is named in a shape, its name can be used in
   this specification to designate either the byte sequence, the
   expression or sequence of expressions it encodes, or the result of
   evaluating those expressions.  When there could be ambiguity, this
   specification specifies which is designated.

2.1.  Parsing algorithm

   The parser operates with a context, which is a list of expressions.
   Each time an expression is parsed, it is appended at the end of the
   context.  The initial context is the abstract yield.

   At the beginning of a BULK stream and after having consumed the byte
   sequence encoding a complete expression, the parser is at the
   dispatch stage.  At this stage, the next byte is a marker byte, which
   tells the parser what kind of expression comes next (the marker byte
   is the first byte of the sequence that encodes an expression).  The
   expression appended to the context after reading a byte sequence is
   called the specific yield of the byte sequence.

   The 0x01 and 0x02 marker bytes are special cases.  When the parser
   reads 0x01, it immediately appends an empty list to the current
   context.  This list becomes the new context.  This new context has
   the previous context as parent.  Then the parser returns to its
   dispatch stage.  When the parser reads 0x02, it appends nothing to
   the context, but instead the parent of the current context becomes
   the new context and the parser returns to the dispatch stage.  Thus
   it is a parsing error to read 0x02 when the context is the abstract
   yield.  It is also a parsing error when the stream ends and the
   parser is not at the dispatch stage in the abstract yield.

   When a form contains at least one expression, the first expression is
   called the operator of that form.  When a form contains more than one
   expression, all but the first expression are called the operands of
   that form.

   Some forms have side-effects in their semantics.  Those side-effects
   MUST NOT affect the parsing of any expression.  They can affect
   evaluation, in which case they MUST only affect the evaluation of
   expressions in the scope of the form.  The scope of an expression is
   the part of its context that follows the expression.  This makes BULK
   lexically scoped.

   Whenever a parsing error is encountered, parsing of the BULK stream
   MUST stop.

   The version of the BULK stream can affect parsing, see Section 4.1.

Thierry                   Expires 31 March 2027                [Page 14]
Internet-Draft                    BULK1                   September 2026

2.1.1.  Summary of marker bytes

                 +========+===================+==========+
                 | marker | shape             |          |
                 +========+===================+==========+
                 | 00     | nil               |          |
                 +--------+-------------------+----------+
                 | 01     | (                 |          |
                 +--------+-------------------+----------+
                 | 02     | )                 |          |
                 +--------+-------------------+----------+
                 | 03     | # Nat {content}   |          |
                 +--------+-------------------+----------+
                 | 04–0F  |                   | reserved |
                 +--------+-------------------+----------+
                 | 10–7F  | Ref               |          |
                 +--------+-------------------+----------+
                 | 80–BF  | w6[value]         |          |
                 +--------+-------------------+----------+
                 | C0–FF  | #[size] {content} |          |
                 +--------+-------------------+----------+

                                  Table 2

2.1.2.  Evaluation

   A processing application MAY implement evaluation of BULK expressions
   and streams.  When evaluating a BULK stream, when the parser gets to
   the dispatch stage and the context is the abstract yield (this is
   called a _dispatch point_), the last expression in the context is
   replaced by what it evaluates to. (of course, this description is
   supposed to provide the semantics of BULK evaluation, but a
   processing application MAY implement evaluation with a different
   algorithm as long as it provides the same semantics)

   Evaluation cannot affect parsing, which means that parsing and
   evaluation can be done in parallel.

   The default evaluation rule is that an expression evaluates to
   itself.  A name within a namespace can have a value, which is what a
   reference associated to this name evaluates to.  A reference whose
   marker value is associated to no namespace or whose name has no value
   evaluates to itself.  How self-evaluating BULK expressions are
   represented in the concrete yield is application-dependent, but
   future specifications MAY define a standard API to access it, similar
   to the Document Object Model for XML.

Thierry                   Expires 31 March 2027                [Page 15]
Internet-Draft                    BULK1                   September 2026

   The evaluation of a form obeys a special rule, though: if the
   operator of the form has type Function, that function is called with
   an argument list and the form evaluates to the return value if it's
   an atom or the evaluation of the return value if it is a form.  If
   the function has type LazyFunction, the argument list is the operands
   of the form.  If the function has type EagerFunction, the argument
   list is the result of evaluating the operands of the form, from left
   to right.  Any expression that has type LazyFunction or EagerFunction
   also has type Function.

   In the abstract yield, if the first expression produced as a value by
   evaluation has type EagerFunction, then after the whole abstract
   yield has been evaluated, the resulting sequence of expressions MAY
   be evaluated as a form.  This is called whole stream evaluation, and
   MAY be a configuration option.  In particular, a processing
   application MAY choose to disable whole stream evaluation during
   streaming (Section 8).

   When this specification describes the evaluation of a form starting
   with a LazyFunction, by default, named shapes designate the
   unevaluated expressions.  When this specification describes the
   evaluation of a form starting with an EagerFunction, by default,
   named shapes designate the result of evaluating expressions.

   A form whose operand doesn't have the type Function evaluates to a
   form containing the result of evaluating each expression of the form,
   from left to right.

   When an application evaluates a BULK expression, it MUST verify that
   evaluation terminates in a finite number of evaluation steps.  An
   application MAY verify finite termination statically or dynamically.
   For example, an application MAY stop evaluation in error after a
   predetermined number of steps.

   Whenever this specification describes the semantics of an expression,
   it describes the result of evaluating that expression, either in
   terms of the value returned or the side-effects executed by
   evaluation of the expression.

   When an evaluation error is encountered, a processing application MAY
   not stop processing with an error.  If the processing application
   produces a partially evaluated concrete yield, it MUST convey enough
   information for the using agent to know the order of all expressions
   in the original BULK stream, which expressions were successfully
   evaluated, and which were not because of evaluation errors.  A
   processing application MAY stop evaluation at the first evaluation
   error and produce the concrete yield in two separated sections, the
   successfully evaluated part, followed by the unevaluated one.  A

Thierry                   Expires 31 March 2027                [Page 16]
Internet-Draft                    BULK1                   September 2026

   processing application MAY continue evaluation and produce the
   concrete yield as a sequence of expressions, each tagged with the
   fact that it was successfully evaluated or not.

   The version of the BULK stream can affect evaluation, see
   Section 4.1.

2.2.  Forms

   starting marker  0x01
      mnemonic: (

   ending marker  0x02
      mnemonic: )

2.2.1.  Difference between sequence and form

   There is a difference between a byte sequence encoding several
   expressions among the current context and a byte sequence encoding a
   form (i.e. a single expression that is a list of expressions).  As an
   example, let's examine several forms of the shape ( foo {bar} ).

   *  In the form ( foo nil nil nil ), {bar} encodes 3 expressions, and
      they are three atoms in the yield.

   *  In the form ( foo nil ), {bar} is a single expression in the
      yield, and that expression is an atom.

   *  In the form ( foo ( nil nil nil ) ), {bar} is also a single
      expression in the yield, and that expression is a form, a list in
      the yield.

   In a shape, when a byte sequence must yield a single expression, it
   has the type Expr.  So the last two examples fit the shape ( foo
   {bar}:Expr ) but not the first.

2.3.  Atoms

2.3.1.  nil

   marker  0x00
      mnemonic: nil

   parsing shape  nil

   Apart from being a possible short marker value, the fact that the
   0x00 byte represents a valid atom means that a sequence of null bytes
   is a valid part of a BULK stream, thus making the format less

Thierry                   Expires 31 March 2027                [Page 17]
Internet-Draft                    BULK1                   September 2026

   fragile.  In a network communication, nil atoms can be sent to keep
   the channel open.  They can also be used as padding at the end of a
   form or between forms.

2.3.2.  Arrays

   Arrays can be used to store arbitrary bytes.

   An array can be interpreted either as a bits sequence or as an
   unsigned integer in binary notation.  The choice depends on the
   context and the application.  Actually, many processing applications
   may not need make any choice, as most programming language
   implementations actually also confuse unsigned integers and bits
   sequences to some extent.  Expressions that are unsigned integers
   (that is, natural numbers) have type Nat (whether they are encoded as
   an array or not).

   Big arrays typically store the content of a file or a binary message
   of another format.  They can also be used to store a vector or matrix
   of fixed-size elements.

   In any case, the semantics of the content must be inferred by the
   processing application; where ambiguity can appear, an application
   SHOULD enclose the array in a form that makes the semantics explicit
   (e.g. string, blob, or unsigned-int).

   Because BULK arrays have no end markers, the payload of a BULK array
   can constitute the end of the stream.

   The start and end of an array are known without reading its content,
   which means that its content can be skipped in constant time and
   mapped in memory (or read lazily by any other means).

   Because BULK can use integers with arbitrary size to store the size
   of an array, BULK arrays have no limit in size.

   Any array also has the type String if its contents can be decoded as
   a string in the current encoding.

   When this specification mentions "the bytes contained in the
   expression {foo}", it never means the bytes encoding that BULK
   expression, but the bytes that are the payload of that expression,
   which could be an array or some other byte container defined in a
   BULK vocabulary.

2.3.2.1.  Generic array

   marker  0x03

Thierry                   Expires 31 March 2027                [Page 18]
Internet-Draft                    BULK1                   September 2026

      mnemonic: #

   parsing shape  # Nat {content}

   After consuming the marker byte, the parser returns to the dispatch
   stage.  It is a parsing error if the parsed expression is not of type
   Nat or if its value cannot be recognized.  This integer is not added
   to any context, but the parser consumes as many bytes as this integer
   and they constitute the content of this array.

   In the text notation, a quoted string is the notation for a generic
   array containing the encoding of that string in the current encoding,
   except if the size of the encoding is below 64 bytes, cf. small
   arrays.

   In the text notation, some text notation enclosed between balanced ([
   and ]) is the notation for a generic array containing the encoding of
   that text notation, except if the size of the encoding is below 64
   bytes, cf. small arrays.

   Types: Bytes, Nat

2.3.2.2.  Small array

   marker  0xC0–0xFF
      mnemonic: #[size]

   parsing shape  #[size] {content}

   The 6 least significant bits of the marker byte are treated as an
   unsigned integer.  This integer is not added to any context, but the
   parser consumes as many bytes as this integer and they constitute the
   content of this array.

   In the text notation, the notation of the marker byte of a small
   array of size X is #[X].  For example, #[2] 0x1234 is a notation for
   the bytes 0xC2-1234.

   In the text notation, a quoted string is the notation for a small
   array containing the encoding of that string in the current encoding
   if the size of the encoding is below 64 bytes.  For example, "abc"
   and #[3] 0x616263 are the notation for the same byte sequence if the
   current encoding is UTF-8.

Thierry                   Expires 31 March 2027                [Page 19]
Internet-Draft                    BULK1                   September 2026

   In the text notation, some text notation enclosed between balanced ([
   and ]) is the notation for a small array containing the encoding of
   that text notation if the size of the encoding is below 64 bytes.
   For example, ([ nil 0 1 256 ]) and #[6] nil w6[0] w6[1] #[2] 0x0100
   are the notation for the same byte sequence.

   Types: Bytes, Nat

2.3.2.3.  Small unsigned integers

   marker  0x80–0xBF
      mnemonic: w6[value]

   parsing shape  w6[value]

   The 6 least significant bits of the marker byte are the value encoded
   by this byte (as bits or as an unsigned integer in binary notation).

   In the text notation, the notation of the marker byte of a small
   unsigned integer of value X is w6[X].  For example, w6[11] is a
   notation for the byte 0x8B (as is 11, cf. Section 4.7).

   Types: Bytes, Nat

2.3.2.4.  Encoding natural numbers

   When the syntax of a BULK form mandates that an expression can only
   be a Nat, an application SHOULD encode it as the smallest possible
   array using one of the following sizes: 6, 8, 16, 32, or any multiple
   of 64 bits.

2.3.3.  Reserved marker bytes

   Marker bytes 0x04−0x0F are reserved for future major versions of
   BULK.  It is a parsing error if a BULK stream with version 1.0
   contains such a marker byte (see Section 4.1).

2.3.4.  References

   marker  0x10−0x7F

   parsing shape  {ns}:1B {name}:1B

      0x7F {ns'} {name}:1B

Thierry                   Expires 31 March 2027                [Page 20]
Internet-Draft                    BULK1                   September 2026

   The {ns} byte is a value associated with a namespace, called the
   namespace marker.  Values 0x10−0x13 are reserved for standard
   namespaces defined by BULK specifications.  Greater values can be
   associated with namespaces identified by a unique identifier.

   The {name} byte is the name index within the namespace.  Vocabularies
   with more than 256 names thus need to be spread across several
   namespaces.

   The specification of a namespace SHOULD include a mnemonic for the
   namespace and for each defined name.  When descriptions use several
   namespaces, the mnemonic of a reference SHOULD be the concatenation
   of the namespace mnemonic, ":" and the name mnemonic if there can be
   an ambiguity.  For example, the fft name in namespace math becomes
   math:fft.

   Type: Ref

2.3.4.1.  Special case

   References have a second parsing rule.  In case a BULK stream needs
   an important number of namespaces, if the marker byte is 0x7F, the
   parser continues to read bytes until it finds a byte different than
   0xFF.  The sum of each of those bytes taken as unsigned integers is
   the namespace marker.  For example, the reference encoded by the
   bytes 0x7F 0xFF 0x8C 0x1A is the name 26 in the namespace associated
   with 522.

3.  Kinds of namespaces

   Standard namespaces have a fixed marker value and are not identified
   by a unique identifier.

   Standard namespaces are immutable.  It is an evaluation error when
   the reference in a name definition is in a standard namespace.

   Extension namespaces are defined with a unique identifier, to be
   associated to a marker value.

   By its decentralized nature, as far as a processing application is
   concerned, while standard namespaces are a special case, there is no
   difference between an extension namespace defined as part of the
   official BULK suite and any other one.

4.  BULK core namespace

   marker  0x10
      namespace mnemonic: bulk

Thierry                   Expires 31 March 2027                [Page 21]
Internet-Draft                    BULK1                   September 2026

       +======+===================================+===============+
       | name | mnemonic                          | type          |
       +======+===================================+===============+
       | 00   | version                           | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 01   | import (namespace, package)       | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 02   | namespace                         |               |
       +------+-----------------------------------+---------------+
       | 03   | package                           |               |
       +------+-----------------------------------+---------------+
       | 04   | define (namespace, package, name) | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 05   | mnemonic                          | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 06   | explain                           | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 07   | string                            | EagerFunction |
       +------+-----------------------------------+---------------+
       | 08   | iana-charset                      | EagerFunction |
       +------+-----------------------------------+---------------+
       | 09   | bulk                              | EagerFunction |
       +------+-----------------------------------+---------------+
       | 0A   | blob                              | EagerFunction |
       +------+-----------------------------------+---------------+
       | 0B   | concat                            | EagerFunction |
       +------+-----------------------------------+---------------+
       | 0C   | indexed-bulk                      | EagerFunction |
       +------+-----------------------------------+---------------+
       | 0D   | indexed-array                     | EagerFunction |
       +------+-----------------------------------+---------------+
       | 0E   | true                              | Boolean       |
       +------+-----------------------------------+---------------+
       | 0F   | false                             | Boolean       |
       +------+-----------------------------------+---------------+
       | 10   | subst                             | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 11   | arg                               |               |
       +------+-----------------------------------+---------------+
       | 12   | rest                              |               |
       +------+-----------------------------------+---------------+
       | 13   | values                            | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 14   | unsigned-int                      | EagerFunction |
       +------+-----------------------------------+---------------+
       | 15   | signed-int                        | EagerFunction |
       +------+-----------------------------------+---------------+
       | 16   | fraction                          | EagerFunction |

Thierry                   Expires 31 March 2027                [Page 22]
Internet-Draft                    BULK1                   September 2026

       +------+-----------------------------------+---------------+
       | 17   | binary-float                      | EagerFunction |
       +------+-----------------------------------+---------------+
       | 18   | decimal-float                     | EagerFunction |
       +------+-----------------------------------+---------------+
       | 19   | prefix                            | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 1A   | postfix                           | LazyFunction  |
       +------+-----------------------------------+---------------+
       | 1B   | arity                             | LazyFunction  |
       +------+-----------------------------------+---------------+

                                 Table 3

4.1.  Version

   shape  ( version {major}:Nat {minor}:Nat )

   When parsing a BULK stream, a processing application MUST determine
   explicitly the major and minor version of the BULK specification that
   the stream obeys.  This information MAY be exchanged out-of-band, if
   BULK is used to exchange a number a very small messages, where
   repeated headers of 6 bytes might become too big an overhead.  A
   processing application MUST NOT assume a default version.

   If the version is expressed within a BULK stream, this form MUST be
   the first in the stream, where its semantics is the side-effect of
   declaring the version.  In any other place, this form evaluates to
   itself.  This specification defines BULK 1.0.  When writing a BULK
   stream's version, an application MUST encode {major} and {minor} by
   the smallest byte sequence as described in Section 2.3.2.4.

   An application writing a BULK stream to long-term storage (e.g. in a
   file or a database record) SHOULD include a version form.

   Two BULK versions with the same major version MUST share the same
   parsing rules and the same definitions of marker bytes used by both.
   Changing the syntax or semantics of existing marker bytes warrants a
   new major version.  Changing the syntax or semantics of existing
   standard names (meaning names in standard namespaces) also warrants a
   new major version.

   It is a parsing error when a processing application encounters a
   major version that it doesn't explicitly support.

   Using marker bytes in the reserved interval, adding standard names,
   or adding new syntactic uses of existing standard names that don't
   overlap with existing uses warrants a new minor version.

Thierry                   Expires 31 March 2027                [Page 23]
Internet-Draft                    BULK1                   September 2026

   If version A and version B of BULK have a different default profile,
   and the change of profile would only affect evaluation of BULK
   streams that use new marker bytes, new standard names or new
   syntactic uses of existing standard names, then the profile
   difference warrants a difference of minor version between A and B.
   Otherwise, the profile change warrants a difference of major version
   between A and B.

   Overall, the goal is that if an application writes a BULK stream of
   version 1.2 that only contains syntactic and semantic elements from
   BULK 1.1, a processing application that only supports BULK 1.1 will
   be able to parse and evaluate that stream.  The writing application
   doesn't need to know the features that are specific to BULK 1.0 and
   BULK 1.1, because as soon as the processing application encounters a
   reserved marker byte, an unknown standard name or an unknown
   syntactic use of a known standard name, it can detect that it cannot
   correctly evaluate the stream.

   For that reason, it is an evaluation error for a processing
   application when there is an unknown syntactic use of a standard
   name, in a BULK stream with a major version supported by the
   processing application but a minor version not supported.

   As far as BULK 1.0 is concerned, using import, define, mnemonic,
   explain, and arity, when none of their operands is either a standard
   name or a form with a standard name as operator, or using version
   somewhere else than as first expression, don't constitute an unknown
   syntactic use.  This makes it possible to overload those names in
   ways that don't interfere with forward-compatibility.

4.2.  Namespaces and packages

4.2.1.  Importing a namespace

   shape  ( import {marker}:Nat ( namespace {id}:Expr ) )

   The semantics of this form is the side-effect of associating the
   namespace identified by {id} to the namespace marker {marker}, within
   the scope of this expression.

   {id} can be a form using one or several names whose namespace marker
   has the value of {marker}. This is called importing a bootstrapping
   namespace, see Section 4.2.4.1.1.

   It is not an evaluation error if the namespace is unknown to the
   processing application.  A processing application MAY produce
   warnings when it encounters this case.

Thierry                   Expires 31 March 2027                [Page 24]
Internet-Draft                    BULK1                   September 2026

4.2.2.  Importing a package

   shape  ( import {base}:Nat ( package {id}:Expr ) {count}:Nat )

   The semantics of this form is the side-effect of associating the
   first {count} namespaces in the package identified by {id} with a
   continuous range of marker bytes starting at {base}, within the scope
   of this expression.

   It is not an evaluation error if the package is unknown to the
   processing application, or if it is known but not all namespaces
   packaged inside are.  A processing application MAY produce warnings
   when it encounters those cases.

   This forms needs an explicit number of namespaces to import to
   preserve transparency: if the number was implicit, there could be
   issues when a processing application that doesn't known how many
   namespaces are in the package wanted to modify the BULK stream.  This
   makes BULK overall simpler and more predictable.

   Example: ( import 21 ( package {foo} ) 3 ) associates the first 3
   namespaces of the package identified by {foo} to the markers 21, 22
   and 23.

   shape  ( import {base}:Nat ( package {id}:Expr ) {count}:Nat
      {increment}:Nat )

   The semantics of this form is the side-effect of associating the
   first {count} namespaces in the package identified by {id} with a
   range of marker bytes starting at {base} at {increment} increments,
   within the scope of this expression.

   Example: ( import 21 ( package {foo} ) 3 2 ) associates the first 3
   namespaces of the package identified by {foo} to the markers 21, 23
   and 25.

4.2.3.  Canonical identifiers

   A processing application MAY use any BULK expression as a namespace
   or package identifier, including atoms like numbers or arrays, but it
   is RECOMMENDED to use one kind of expression called a _canonical
   identifier_. Canonical identifiers have the shape ( {type}:Ref
   {content}:Bytes ).  The role of {type} is to describe how to
   interpret {content}: is it a URI, a UUID, a SHA-3 checksum?

Thierry                   Expires 31 March 2027                [Page 25]
Internet-Draft                    BULK1                   September 2026

   Canonical identifiers can be used to efficiently encode IPLD's
   Content IDentifiers in BULK, but they have a broader use.  Where
   IPLD's CIDs can only encode content-adressing, canonical identifiers
   can encode arbitrary identifiers, including some that can be created
   independently of the content, like UUIDs.

4.2.4.  Namespace definition

   shape  ( define ( namespace {id}:Expr ) {def}:Bytes )

   The semantics of this form is the side-effect of defining a new
   namespace.  The bytes contained in {def} MUST be a BULK stream with a
   version form, and have the shape ( version {major}:Nat {minor}:Nat )
   {config}:Expr {definitions}. {config} MUST be a form and its first
   element MUST have the shape ( namespace {marker}:Nat ).  The state of
   the namespace associated to {marker} after evaluating all expressions
   in {definitions} (starting with a namespace with no definitions, no
   mnemonics, and no documentations) is made the definition of the
   namespace identified by {id}, within the scope of this expression.
   This also associates that namespace to the namespace marker {marker},
   within the scope of this expression.

4.2.4.1.  Verifiable namespace definition

   When a processing application recognizes that {id} designates a
   digest that matches the bytes contained in {def}, this creates a
   verifiable namespace.

   If more data than {id} is needed to verify {id} against the bytes
   contained in {def} (like the salt of a hash function, or the
   namespace of a UUID), this data MUST be provided in {config}.

   It is an evaluation error if the processing application recognizes
   that {id} designates a digest but it doesn't match the bytes
   contained in {def}.

   Verifiable namespaces are meant to be immutable, but that would be
   circumvented if they were built upon namespaces that aren't.  A
   verifiable namespace that only uses names from immutable namespaces
   is an immutable namespace (see Section 3).

   A processing application SHOULD only consider digest algorithms that
   are currently known to be cryptographically secure for the
   determination of namespace and package immutability.  For example, a
   processing application could check that a namespace with an MD5
   identifier is verifiable, but it SHOULD NOT determine it to be
   immutable.

Thierry                   Expires 31 March 2027                [Page 26]
Internet-Draft                    BULK1                   September 2026

   The BULK stream in {def} has a version form to prevent the
   possibility that a verified namespace definition could be evaluated
   to different results by applications using different BULK versions.
   The definitions are in a nested BULK stream because if the definition
   could use immutable namespaces imported outside of the definition,
   the same verifiable definition could be used in different contexts,
   also defeating the immutability.

4.2.4.1.1.  Bootstrapping verification

   When using canonical identifiers (see Section 4.2.3), a verifiable
   namespace will use a form to express its digest.  For immutable
   namespaces to exist, this means that at least one namespace needs to
   express its own digest with a name from within itself.  This is
   called a bootstrapping namespace.  A bootstrapping namespace MUST use
   a canonical identifier.

   When importing a bootstrapping namespace, the processing application
   looks up in its known namespaces if there is a namespace identified
   by a form whose operator is a name in itself with the same name index
   as the operator in the identifier in the import.  If that name is
   associated with a digest algorithm and the logic of the digest
   algorithm determines that the digest form in the import matches the
   digest in the namespace definition (along with additional
   configuration data provided there), then the bootstrapping is
   successful (meaning that the definition has been found and can be
   imported).

   This process doesn't rely just on the digest in the import being
   identical to the digest in the known definition, because some digests
   algorithms (e.g. extended-output functions) have been designed to
   retain some collision resistance when using a prefix of the digest,
   so a definition could contain a 256 or 512 bits wide digest, but some
   small BULK streams could import it with the first 64 or 128 bits,
   when it is worth the trade-off between stream size and collision
   resistance.

   It is not an evaluation error if bootstrapping fails.  The only
   consequence is that the bootstrapping namespace is not known, and any
   other namespace using this namespace for its identifier will not be
   known either.  A processing application MAY produce warnings when it
   encounters this case.

   The same logic can be used to import a bootstrapping package, which
   is a package identified with a form from one of its own namespaces.

Thierry                   Expires 31 March 2027                [Page 27]
Internet-Draft                    BULK1                   September 2026

   When using immutable namespaces, bootstrapping packages are likely to
   be norm, as using almost any immutable namespace requires importing
   the bootstrapping namespace used to identify the target namespace,
   then that namespace.  Every definition of an immutable namespace can
   be accompanied by the definition of a bootstrapping package packaging
   that namespace, its bootstrapping namespace and any other namespaces
   that are likely to be used with it.

4.2.5.  Package definition

   shape  ( define ( package {id}:Expr ) {def}:Bytes )

   The semantics of this form is the side-effect of creating a package
   identified by {id}. The bytes contained in {def} MUST be a BULK
   stream with a version form and have the shape ( version {major}:Nat
   {minor}:Nat ) {config}:Expr {preamble} {namespaces}:Expr. {config}
   MUST be a form. {namespaces} MUST be a form containing a sequence of
   expressions each identifying a BULK namespace.

   When a processing application recognizes that {id} designates a
   digest that matches the bytes contained in {def}, this creates a
   verifiable package.

   If more data than {id} is needed to verify {id} against the bytes
   contained in {def} (like the salt of a hash function, or the
   namespace of a UUID), this data MUST be provided in {config}.

   It is an evaluation error if the processing application recognizes
   that {id} designates a digest but it doesn't match the bytes
   contained in {def}.

   Packages are meant to be immutable, but that would be circumvented if
   they were built upon namespaces that aren't.  A verifiable package
   that only contains immutable namespaces is an immutable package.

4.2.6.  Name definitions

   To define a reference is to make a value the semantics of any
   reference with the same associated namespace and the same name, in
   the scope of that definition.

   This change of semantics operates on the namespace as identified by
   its unique identifier, not the marker value.  This means that if a
   namespace Foo is associated to markers 21 and 22 and a reference with
   namespace marker 21 and name 0 is defined to true, then in the scope
   of that definition, a reference with namespace marker 22 ans name 0
   will be evaluated as true.

Thierry                   Expires 31 March 2027                [Page 28]
Internet-Draft                    BULK1                   September 2026

   When a BULK stream containing definitions for a namespace comes from
   a trusted source (i.e. in configuration files of the application, or
   in the communication with an agent that has been granted the relevant
   authority), an application MAY give those definitions long-lasting
   semantics (i.e. keep the values of the names at the end of parsing).
   This is the RECOMMENDED mechanism for bulk namespace definition when
   the semantics of the defined expressions can be expressed completely
   by BULK expressions (see Appendix B).

   shape  ( define {ref}:Ref {value}:Expr )

   The semantics of this form is the side-effect of defining the
   reference {ref} to the value of evaluating {value}.

4.2.7.  Mnemonic

   shape  ( mnemonic ( namespace {marker}:Nat ) {mnemonic}:Expr )

   This shape declares the value of evaluating {mnemonic} to be the
   mnemonic for the namespace associated with the marker {marker}.

   shape  ( mnemonic Ref {mnemonic}:Expr )

   This shape declares the value of evaluating {mnemonic} to be the
   mnemonic for the name designated by the reference.

4.2.8.  Explain

   shape  ( explain ( namespace {marker}:Nat ) {doc}:Expr )

   This shape declares the value of evaluating {doc} to be the
   documentation for the namespace associated with the marker {marker}.

   shape  ( explain Ref {doc}:Expr )

   This shape declares the value of evaluating {doc} to be the
   documentation for the name designated by the reference.

   Documentation expressions can be strings with plain text or use any
   BULK vocabulary to use a richer documentation format (including BULK
   forms to make a format of plain text explicit).

4.3.  Strings and other typed byte arrays

4.3.1.  Current encoding

   shape  ( define string {encoding}:Expr )

Thierry                   Expires 31 March 2027                [Page 29]
Internet-Draft                    BULK1                   September 2026

   The semantics of this form is the side-effect that, in the scope of
   this expression, the default encoding for expressions that are
   understood by the application as character strings is the encoding
   designated by {encoding}.

   As the abstract yield doesn't contain strings but expressions that
   will be used as strings by the application, it is not a parsing error
   if the application doesn't recognize {encoding}. In this situation,
   it is only a parsing error when the application actually needs to
   decode a byte sequence as a string with that encoding.  It is not a
   parsing error when a processing application only transmits a byte
   sequence encoding a string, if it can accurately convey the encoding
   to the receiving application.

4.3.2.  String

   shape  ( string {string}:Expr )

   This form indicates that the bytes contained in the expression
   {string} are meant to be interpreted as a string encoded with the
   current default string encoding.

4.3.3.  String with explicit encoding

   shape  ( string {encoding}:Expr {string}:Epxr )

   This form indicates that the bytes contained in the expression
   {string} are meant to be interpreted as a string encoded with the
   encoding designated by {encoding}.

4.3.4.  IANA registered character set

   shape  ( iana-charset {id}:Expr )

   This designates the string encoding registered among the IANA
   Character Sets [IANA-Charsets] whose MIBenum is {id}.

   Type: Encoding.

4.3.5.  Nested BULK stream

   shape  ( bulk {bulk}:Expr )

   This form indicates that the bytes contained in the expression {bulk}
   are meant to be interpreted as a BULK stream.  If the stream doesn't
   start with a version form, the stream explicitly has the same version
   as the parent stream.

Thierry                   Expires 31 March 2027                [Page 30]
Internet-Draft                    BULK1                   September 2026

   The semantics of this form is the same as the evaluation of the BULK
   stream in {bulk} taken as a form.  For example, these two forms have
   the same evaluation:

   *  ( 4 5 )

   *  ( bulk true #[2] 4 5 )

   This form can be useful to let the application reading a BULK stream
   skip parsing a large section.  In that case, it MUST be enclosed in a
   values form to prevent the security issue described in Section 9.4.

4.3.6.  Blob

   shape  ( blob {blob}:Expr )

   This form indicates that the bytes contained in {blob} are meant be
   interpreted as just a raw sequence of bytes, not to be decoded.

4.4.  Array operations

4.4.1.  Array concatenation

   shape  ( concat {arrays} )

   The value of this form is an array that contains the bytes contained
   in the first expression in {arrays} followed by the bytes contained
   in the second expression, and so on for each expression.  It is an
   evaluation error if any expression in {arrays} doesn't evaluate to a
   container of bytes.

   ( concat ) evaluates to an empty array like #[0].

4.4.2.  Indexed data

   When writing a stream containing a big number of expressions where an
   application may want to access one of those expression without
   parsing all expressions before, one could imagine as a solution to
   use pointer-like references that each use the offset of some
   expression in the stream.  This solution creates a security risk,
   because if reading according to the pointers doesn't produce the same
   result as parsing the stream without using them, an attacker might
   use this inconsistency to their advantage, when they can expect one
   application to use pointers and another application to use normal
   parsing, especially when the stream is big enough that verifying the
   consistency of the pointers might be costly enough that it might not
   be done or not in time to prevent the attack.

Thierry                   Expires 31 March 2027                [Page 31]
Internet-Draft                    BULK1                   September 2026

   Because of that risk, whenever a stream includes indexed BULK
   expressions, that is, expressions that are meant to be accessed by
   their byte position, indexed reading SHOULD be the only way used to
   access them.  To that end, indexed data SHOULD be stored in arrays.

   When the goal of indexed data is to selectively parse only part of
   the BULK stream, a values form MUST be used to prevent the security
   issue described in Section 9.4.

4.4.2.1.  Indexed BULK expression

   shape  ( indexed-bulk {container}:Expr {start}:Expr )

   The semantics of this form is the value of evaluating the BULK
   expression starting at offset {start} in the bytes contained in
   expression {container}.

   Beware that, although evaluations of all other BULK definitions
   follow lexical scoping, any definition used inside an indexed
   expression that isn't defined inside that same indexed expression
   follows dynamic scoping with respect to any places where it's used.
   Any definition made inside an indexed expression still follows
   lexical scoping.

4.4.2.2.  Indexed array

   shape  ( indexed-array {container}:Expr {start}:Expr {size}:Expr )

   The semantics of this form is the value of an array whose content are
   {size} bytes, starting at offset {start} in the bytes contained in
   the expression {container}.

   shape  ( indexed-array {container}:Expr {start}:Expr )

   The semantics of this form is the value of an array whose content are
   the bytes starting at offset {start} in the array {container} until
   its end.

   Compared to indexed-bulk, which can reference an array expression,
   indexed-array is useful when several different but overlapping
   sections of the same byte sequence are needed as arrays, or when
   reversing packing (Section 5.1) through evaluation (to avoid packing
   the marker bytes).

4.5.  Booleans

   shape  true

Thierry                   Expires 31 March 2027                [Page 32]
Internet-Draft                    BULK1                   September 2026

   shape  false

   Type: Boolean.

4.6.  Substituton

4.6.1.  Substitution function

   shape  ( subst {code} )

   The semantics of this form is an expression of type LazyFunction,
   called the substitution function.  The semantics of the substitution
   function are the semantics of {code}, but where the names arg and
   rest are substituted as such:

   *  ( arg {n}:Nat ) is replaced by the element number {n} (starting at
      zero) of the substitution function's arguments list.

   *  ( rest {n}:Nat ) is replaced by the substitution function's
      arguments list without its first {n} elements.

   It is an evaluation error if the substitution function is called with
   too few arguments with respect to the arg and rest forms in {code}.

4.6.2.  Examples

   Here is a definition of the inverse followed by the numbers 1/2, 1/3
   and 1/4:

   ( define inverse ( subst ( fraction 1 ( arg 0 ) ) ) )
   ( inverse 2 )
   ( inverse 3 )
   ( inverse 4 )

   Substitution will splice multiple expressions in place:

   The evaluation of:

   ( define foo ( subst 20 ( rest 0 ) 50 ) )
   ( 10 ( foo 30 40 ) 60 )

   must produce: ( 10 20 30 40 50 60 )

4.6.3.  Isolated scope

   shape  ( values {exprs} )

Thierry                   Expires 31 March 2027                [Page 33]
Internet-Draft                    BULK1                   September 2026

   This form's semantics is the values produced by the evaluation of
   {exprs}. This means that the side-effects in {exprs} can affect its
   own evaluation, but not the scope of the values form.  It makes it
   possible to isolate evaluation side-effects.

   This MUST be used for selectively parseable data, to prevent the
   security issue described in Section 9.4.  Because the semantics of
   this form is only the values produced by the evaluation of its
   operands, this evaluation can wait until those values are actually
   needed.

   This property means that if a processing application uses lazy
   evaluation of BULK expressions, every expression in a values form is
   selectively evaluated only when needed, which in turns means that any
   nested BULK stream within a values form is selectively parsed only
   when needed.

4.7.  Arithmetic

   A processing application must recognize the type of all expressions
   defined in this specification that have the type Nat, but an
   application MAY consider a number as having an unknown value if it
   can't decode its value or has no adequate data type to store it.  It
   is only a parsing error if the number is needed by the parsing
   algorithm.  It is only an evaluation error if the number is needed by
   the evaluation algorithm.

   In the text notation of a BULK stream, a decimal integer is the
   notation for the smallest byte sequence that yields this integer as
   described in Section 2.3.2.4.  For example, ( 31 256 ) is a notation
   for the bytes 0x01 0x9F 0xC2-0100 0x02.

4.7.1.  Unsigned integer

   shape  ( unsigned-int {bits}:Expr )

   The semantics of this form is the value of the unsigned integer
   represented in binary notation in the bits contained in {bits}. This
   form exists in case disambiguation of the semantics of an array (or
   another bit container) is necessary.

   Type: Number, Real, Int, Nat.

4.7.2.  Signed integer

   shape  ( signed-int {bits}:Expr )

Thierry                   Expires 31 March 2027                [Page 34]
Internet-Draft                    BULK1                   September 2026

   The semantics of this form is the value of the signed integer
   represented in two's-complement notation in the bits contained in
   {bits}.

   Type: Number, Real, Int.

4.7.3.  Fraction

   shape  ( fraction {num}:Expr {div}:Expr )

   The semantics of this form is the fraction with denominator {num} and
   divisor {div}.

   Type: Number.

4.7.3.1.  Fixed-point numbers

   fraction makes it possible to express fixed-point numbers in BULK:

   *  ( fraction 15 4 ) has value 11.11_2 (3.75_10)

   *  ( fraction 123 100 ) has value 1.23

   A more extensive arithmetic vocabulary could define forms to express
   fixed-point numbers according to a given base, as well as predefined
   fixed points (e.g. the use of fixed-point numbers with two decimals
   is pretty common with financial data).

4.7.4.  Binary floating-point number

   shape  ( binary-float {bits}:Expr )

   The semantics of this form is the floating-point number expressed in
   IEEE 754-2008 binary interchange format by the bits contained in
   {bits}. {bits} can be of size 16, 32, 64, 128 or any bigger multiple
   of 32 bits, as per IEEE 754-2008 rules.

   Types: Number, Real, Float.

4.7.5.  Decimal floating-point number

   shape  ( decimal-float {bits}:Expr )

   The semantics of this form is the floating-point number expressed in
   IEEE 754-2008 decimal interchange format by the bits contained in
   {bits}. {bits} can be of size 32, 64, 128 or any bigger multiple of
   32 bits, as per IEEE 754-2008 rules.

Thierry                   Expires 31 March 2027                [Page 35]
Internet-Draft                    BULK1                   September 2026

   Types: Number, Real, Float.

4.8.  Bytecodes

   This specification and other official BULK specifications use forms
   with a reference operator as their basic building blocks.  Basically,
   these are a binary representation of an abstract syntax tree.  As
   noted previously, this means that most representations weigh 4 bytes
   plus their actual content, which will in turn have some overhead
   because of one or several marker bytes.

   But when there is a special need for compactness, BULK makes it
   possible to design protocols and formats with different trade-offs,
   while retaining its property of being parseable by processing
   applications not knowing the protocol in its entirety.

   On one end of the spectrum, a format might choose to use an array to
   encapsulate an ad hoc binary format.  An extreme use of this scheme
   would be to use BULK just to make explicit the binary format used and
   for nothing else.  With a known profile (Section 6) (for example with
   a file extension and/or media type for such explicitly typed BLOBs),
   such a BULK stream can consist solely of the version form, a
   reference that describes the binary format and an array, which would
   amount to an overhead between 11 bytes and 20 bytes depending on the
   size of the content (11, 13, 14, 16 and 20 bytes for contents of no
   more than 63B, 255B, 65kB, 4GB and 18EB respectively).  Without a
   profile, with the namespaces associations in a package, the minimum
   overhead is only between 32 and 41 bytes (the difference is a single
   import form, assuming a digest of 64 bits).

   Still, even this extreme in the design space retains the ability to
   insert expressions in the BULK stream, whatever their type.  Thus
   metadata can be added about data that is represented in a format that
   doesn't allow for metadata or that allows only for limited metadata.
   Appendix F gives a few examples of what encapsulating existing media
   types in BULK could bring.

   In-between these two extremes, several options are available to
   produce a format that leverages the BULK parser a lot more while
   being more compact than a basic BULK format.  The following forms
   provide a standard way to create such formats, called BULK bytecodes.

   A BULK bytecode is a flat sequence of expressions.  The evaluation of
   a bytecode form transforms that sequence to an abstract syntax tree
   of its contents (and then the resulting expression can be evaluated
   with the normal BULK evaluation rules).  The expressions of the
   bytecode are divided among bytecode operators and bytecode operands.
   Operators are references that will end up as form operators in the

Thierry                   Expires 31 March 2027                [Page 36]
Internet-Draft                    BULK1                   September 2026

   abstract syntax tree.  Operands are all other expressions.  Prefix
   bytecodes are those where operators come before their operands,
   postfix bytecodes are those where operators come after their
   operands.  In the following forms, operators MUST be references.

   When evaluating a bytecode, it is an evaluation error when the
   processing application encounters a reference for which it cannot
   determine if it is an operator or its arity (the number of operands
   it will have).  An expression that is not a reference is always an
   operand.

4.8.1.  Prefix bytecode

   shape  ( prefix {bytecode} )

   This is a prefix bytecode form.  The bytecode to be transformed is
   the sequence of expressions in {bytecode}.

   To transform a prefix bytecode, a processing application creates an
   alternate context.  If the first expression of the bytecode is an
   operand, it is removed from the beginning of the bytecode and
   appended at the end of the alternate context.  If the first
   expression of the bytecode is an operator, it is removed from the
   beginning of the bytecode and a list is created with the operator as
   the first expression, then as many next expressions as its arity are
   removed from the beginning of the bytecode and appended at the end of
   this list.  Then that resulting list is appended at the end of the
   alternate context.  The transformation continues until the bytecode
   is empty, in which case the transformation is complete and the
   alternate context is the value of evaluating the bytecode form.  The
   resulting form can then be evaluated in turn.

   Example: the evaluation of

   ( define ( arity prefix )
     ( nil game ) ( 2 black ) )
   ( prefix game black 1 2 black 3 4 black 5 6 )

   is

   ( game
    ( black 1 2 )
    ( black 3 4 )
    ( black 5 6 ) )

   It is an evaluation error when there are less expressions remaining
   in the bytecode than the arity of the current operator.  In the error
   information, a processing application MAY provide the alternate

Thierry                   Expires 31 March 2027                [Page 37]
Internet-Draft                    BULK1                   September 2026

   context and the remaining bytecode.  The alternate context after a
   failed transformation MUST NOT appear in the abstract yield as if
   evaluation had successfully transformed the bytecode.

4.8.2.  Postfix bytecode

   shape  ( postfix {bytecode} )

   This is a postfix bytecode form.  The bytecode to be transformed is
   the sequence of expressions in {bytecode}.

   To transform a postfix bytecode, a processing application creates a
   data stack.  If the first expression of the bytecode is an operand,
   it is removed from the beginning of the bytecode and pushed on top of
   the stack.  If the first expression of the bytecode is an operator,
   it is removed from the beginning of the bytecode and a list is
   created with the operator as the first expression, then as many next
   expressions as its arity are popped from the stack and appended at
   the end of this list (with the top of the stack as the last element).
   Then that resulting list is pushed on top of the stack.  The
   transformation continues until the bytecode is empty, in which case
   the transformation is complete and the list of expressions on the
   stack (with the top of the stack as the last element) is the value of
   evaluating the bytecode form.  The resulting form can then be
   evaluated in turn.

   Example: the evaluation of

   ( define ( arity postfix )
    ( nil game ) ( 2 black white comment alternative ) )
   ( postfix
     game
     1 2 black
     "white tried an unorthodox opening" 3 4 white comment
     "a more classical opening would be" 8 9 white comment
     alternative
     2 3 black
     4 5 white )

   is

   ( game
     ( black 1 2 )
     ( alternative
       ( comment "white tried an unorthodox opening" ( white 3 4 ) )
       ( comment "a more classical opening would be" ( white 8 9 ) ) )
     ( black 2 3 )
     ( white 4 5 ) )

Thierry                   Expires 31 March 2027                [Page 38]
Internet-Draft                    BULK1                   September 2026

   The obvious advantage of postfix bytecode is that it makes it
   possible to compact nested forms when they have a known arity.  When
   a reference in a vocabulary can be used in a form containing a
   variable number of expressions, if some arity is used frequently
   enough, an application can define a specific form for it.  The trade-
   offs for this are explained in Appendix D

   It is an evaluation error when there are less expressions remaining
   on the data stack than the arity of the current operator.  In the
   error information, a processing application MAY provide the data
   stack and the remaining bytecode.  The data stack after a failed
   transformation MUST NOT appear in the abstract yield as if evaluation
   had successfully transformed the bytecode.

4.8.3.  Arity definition

   shape  ( define ( arity {contexts} ) {arities} )

   This form defines the arity of references in the context of
   bytecodes.

   {contexts} can contain prefix, postfix, or any other reference, to
   specify in the context of which kind of bytecode the arities are
   modified.  If {contexts} is empty, the arities are modified in the
   context of all kinds of bytecodes.

   {arities} is a sequence of expressions that each can be shaped as
   such:

   *  nil: meaning all known arities should be forgotten

   *  ( {kind}:Expr {target} ):

      -  if {kind} is nil, it sets all references designated by {target}
         as operands

      -  if {kind} is typed Nat, it sets all references designated by
         {target} as operators of arity {kind}

      -  if {target} is nil, it designates all references with unknown
         arity

      -  if {target} is a sequence of references, it designates each of
         those

Thierry                   Expires 31 March 2027                [Page 39]
Internet-Draft                    BULK1                   September 2026

5.  Optimizing compactness

5.1.  Packing

   If the overhead of several marker bytes in some operands is too much,
   more compactness can be achieved by packing together small operands.
   For example, instead of an operator with two integers as its
   operands, one could specify an operator to take a single array as
   operand and extract the integers from it.  When the processing
   application does this extraction, the format retains the ability to
   operate on many sizes of integers, because the processing application
   can still deduce the size of the integers by dividing the size of the
   array by two.  This can be used outside or inside of bytecodes (to
   stack the compacting effects of both).

   For example, a BULK format representing player moves with a pair of
   coordinates on a large board might represent a single move with the
   following shapes:

   basic (8 bytes)  ( move/2 #[1] 0x41 #[1] 0x5A )

   packed basic (7 bytes)  ( move/1 #[2] 0x41 0x5A )

   bytecode (6 bytes)  move/2 #[1] 0x41 #[1] 0x5A

   packed bytecode (5 bytes)  move/1 #[2] 0x41 0x5A

   Packing can also be done without adding its burden on the logic of
   the processing application, by using evaluation to transform packed
   forms into simpler forms (but they need to be created separately for
   each size of operands).  For example, the following:

   ( define move/1-16
     ( subst ( move/2 ( indexed-array ( arg 0 ) 0 1 )
                      ( indexed-array ( arg 0 ) 1 1 ) ) ) )
   ( move/1-16 #[2] 0x415A )

   Would be evaluated into:

   ( move/2 #[1] 0x41 #[1] 0x5A )

   More complex packing can be encoded as well.  For example, this would
   be a form that packs two 16-bits unsigned integers, one 32-bits
   signed integer, one reference, and one variable-length string:

Thierry                   Expires 31 March 2027                [Page 40]
Internet-Draft                    BULK1                   September 2026

   ( define pack-224-ref-str
     ( subst ( bar ( indexed-array ( arg 0 ) 0 2 )
                       ( indexed-array ( arg 0 ) 2 2 )
                       ( signed-int ( indexed-array ( arg 0 ) 4 4 ) )
                       ( indexed-bulk ( arg 0 ) 8 )
                       ( indexed-array ( arg 0 ) 10 ) ) )

   In essence, packing makes it possible to embed very simple ad hoc
   binary formats described within the BULK framework.

5.2.  Mixing literals

   The transformation defined for the bytecode forms makes it possible
   to mix literal expressions and operations represented by a sequence
   of operators and operands.  A typical example would be that instead
   of explicitly encoding which player is playing in turn, to only
   encode the move and let the player information be implicit, when the
   order of players is dictated by the game logic.  In the previous go
   example, for instance, one might represent each alternating move by
   the two players as two integers, lowering the weight of each normal
   move to 2 bytes as coordinates are below 64:

   ( define ( arity postfix )
    ( nil game ) ( 2 black white comment alternative ) )
   ( postfix
     game
     1 2
     "white tried an unorthodox opening" 3 4 white comment
     "a more classical opening would be" 8 9 white comment
     alternative
     2 3
     4 5 )

   The difference between all these schemes and an array containing
   fixed-size elements is that you keep the ability to insert other
   forms, like here to represent comments on the game or variants.

5.3.  Trade-offs

   The most visible cost of the bytecode format is that if it contains
   operators whose arity is unknown to a processing application, the
   whole list after the first occurrence of them is unreadable to that
   processing application, whereas in the basic format, the processing
   application can still process all the forms it understands, and that
   requires no anticipation by the application creating the BULK stream.

Thierry                   Expires 31 March 2027                [Page 41]
Internet-Draft                    BULK1                   September 2026

   The only case where operators could have an unknown arity is when the
   application writing the stream didn't include the arities of every
   operator used in the stream to avoid the redundancy with their
   previous definition (typically in the definition of their respective
   namespaces).  That redundancy would be offset by the space reduction
   of postfix bytecode for streams containing a few dozens forms.  At
   that point, with all arities explicit, with packing and with literals
   for the most used forms, postfix bytecode gets on the Pareto front
   for size and generality, retaining the full generality of BULK while
   saving a lot of space.

   It is RECOMMENDED that whenever explicit arities would be a small
   fraction of the total stream size, they all be given.  A processing
   application MAY choose to never include explicit arities for the
   names of the main namespace of the format, if that namespace's
   definition includes arities, when processing the stream without
   knowledge of that definition wouldn't make sense.

   But another cost of the bytecode format is the loss of resilience.
   If a BULK stream ends up with errors in the content, whether during
   creation or transit, but those errors don't affect the syntactic
   structure of the stream, then those errors will only prevent
   processing the form they are in.  If errors occur within a metadata
   form, the data is still readable.  If errors occur within one entry
   in an archive, other entries are still readable.  But a far wider
   class or errors will make a bytecode impossible to evaluate, and the
   blast radius is everything after the error.

   When a protocol needs the exchange of messages where every byte
   spared counts and there is no sense in trying to recover a partial
   message after corruption, then using a packed bytecode is probably an
   excellent solution.  When writing large quantities of small data
   elements to long-term storage, when overhead adds up significantly,
   it is RECOMMENDED to use a mechanism to limit the blast radius of
   possible data corruption.

   There are several possible solutions to limit where bytecode errors
   propagate: one is chunking the bytecode into several bytecode forms,
   another is inserting at regular intervals a beacon expression that
   can never be an operand, and that can "reset" the bytecode
   transformation process that had been corrupted before.  The former is
   simpler but the latter might be better suited when data is streamed
   (see Section 8).

Thierry                   Expires 31 March 2027                [Page 42]
Internet-Draft                    BULK1                   September 2026

6.  Profiles

   A profile is a byte sequence parsed by a processing application just
   after the version form or before the first expression if there is no
   version form.  Thus a parser SHOULD look ahead at the beginning of a
   stream to see if the first three bytes are ( bulk:version.  With
   respect to the BULK stream, the profile is an out-of-band
   information, usually implicit.

   A processing application doesn't need to actually parse the profile
   or include the profile's yield in the concrete yield, as long as the
   semantics of the abstract yield are maintained.

   The same BULK stream might be processed with different profiles.

   A processing application MUST NOT deduce the profile from the content
   of a BULK stream.

6.1.  Profile redundancy

   A processing application SHOULD only rely on the use of a profile
   when it is a safe assumption that the profile is known, for example
   within a communication where the protocol dictates the profile.

   In particular, long-term storage of a BULK stream SHOULD preserve
   profile information, for example with a media type that dictates the
   profile.

   Otherwise, an application writing a BULK stream in a long-term
   storage SHOULD include the profile after the version form.  For this
   reason, the expressions in a profile SHOULD have idempotent
   semantics.

6.2.  Standard profile

   This specification defines the default profile that a processing
   application MUST use when it is not using a specific profile:

   ( define string ( iana-charset 106 ) )

   This means that the default string encoding in a BULK stream is UTF-
   8.

Thierry                   Expires 31 March 2027                [Page 43]
Internet-Draft                    BULK1                   September 2026

6.3.  Fixed BULK: all profile, no evaluation

   Fixed BULK is a mode of processing for a format or protocol where
   evaluation has been deemed detrimental (it could be that it's too
   expensive computationally, or that it adds too much complexity in the
   processing application's code or the protocol).  In Fixed BULK mode,
   a processing application uses a profile that contains one or several
   namespace associations and possibly definitions.  Because evaluation
   is disabled, no namespace association or definition will be executed
   during processing and namespaces are "fixed" to their markers as per
   the profile.

   Fixed BULK mode lets a format or protocol use BULK's syntax while
   operating more like binary format frameworks, like ASN.1, Protocol
   Buffers, or CBOR.  An interesting difference is that the
   concatenation of the profile and the Fixed BULK stream is a normal
   BULK stream.

7.  Discovery of namespaces and packages

   When a processing application encounters an unknown namespace or
   package identifier, it MAY ask several sources to provide the
   definition.  This is called _discovery_. The possible sources include
   the agent that made use of the unknown identifier, known BULK
   registries (registries of definitions for BULK namespaces or
   packages), or protocols based on content-addressing.  When discovery
   is done without user intervention while processing a BULK stream, in
   order to evaluate it fully, it is called _immediate discovery_. When
   missing identifiers are collected to be retrieved later, under human
   supervision, it is called _deferred discovery_.

   The fundamental security risk in this mechanism is when an attacker
   manages to be the first to use some identifier and poison the
   processing application by giving it a malicious definition, or
   managed to poison one or several BULK registries.  To avoid that
   risk, a processing application MUST only accept definitions for
   immutable namespaces and immutable packages during immediate
   discovery.  It is also RECOMMENDED that a BULK registry only store
   definitions for immutable namespaces and immutable packages.

   During deferred discovery, a processing application MUST NOT accept
   namespace or package definitions from an untrusted source when they
   are not immutable (including when the digest doesn't match the data,
   or when the identifier form is not known to be a digest by the
   processing application).  A BULK registry MUST NOT store namespace or
   package definitions from an untrusted source when they are not
   immutable.

Thierry                   Expires 31 March 2027                [Page 44]
Internet-Draft                    BULK1                   September 2026

   While discovery of bootstrapping namespaces and packages can be done,
   the digest algorithm used in the identifier of a bootstrapping
   namespace or package MUST have been known before discovery (this can
   only be checked after discovery has retrieved and evaluated the
   definitions).  This is possible when a bootstrapping package uses a
   digest from a namespace that was known by the processing application
   before discovery, or when the digest name comes from a namespace
   where it is aliased to a digest name that was known by the processing
   application before discovery.

   With immutable namespaces and immutable packages, though, automatic
   discovery can be a safe mechanism if some risks are mitigated:

   *  Asking untrusted registries might expose when other agents
      communicate with the processing application.  Making the request
      asynchronously, with noise in the timing, can mitigate that risk.
      Application operators need to consider the trade-off between
      latency and privacy, or give the option to agents to make that
      determination.  Making the request in ways that hide the
      processing application's identity, like using The Onion Router, is
      another option.

   *  Asking untrusted registries might expose what kind of data is sent
      by other agents to the processing application.  Making the request
      in ways that hide the processing application's identity, like
      using The Onion Router, is an option.

   *  Asking the agent or not might expose the fact that a namespace was
      already known or not.  Some processing applications SHOULD provide
      configurable policies so that operators can choose one that is
      relevant to the sensitivity of the services they operate.
      Examples could include: always asking for namespaces that haven't
      been explicitly marked safe, always or never asking to some
      classes of agents.

8.  Streaming

   The parsing and evaluation algorithms allow for reading a BULK stream
   while it has not been completely received by the processing
   application.  Once the parser reaches a dispatch point (see
   Section 2.1.2), the fully parsed expression can be evaluated if
   evaluation is enabled.

   This makes BULK usable for streaming data, including in the case of a
   communication protocol with a long-lived connection where requests
   and responses must be processed immediately, including a full-duplex
   communication (see Section 8.2).

Thierry                   Expires 31 March 2027                [Page 45]
Internet-Draft                    BULK1                   September 2026

8.1.  Broadcasting BULK

   It is possible to stream BULK data in a broadcast setting, meaning
   that a processing application could start receiving data mid-stream
   and never see the data sent before it started listening to the
   broadcasted stream.

   A broadcasted BULK stream MUST NOT contain expressions whose
   evaluation have side-effects when their scope is the abstract yield.
   If it contained such expressions, the same processing application
   that started listening at different points in the stream could
   produce different results for a common portion of the concrete yield.

   This means that a protocol that employs BULK broadcasting and uses
   any extension namespace MUST provide a profile, either in the
   protocol specification, or during connection establishment.  For the
   latter, when a BULK stream is broadcasted by HTTP, the server can use
   the bulk-profile link relation in headers (see Section 10.2).

   When broadcasting BULK, one issue is stream capture, to discover an
   offset in the stream that is a dispatch point.  This specification
   describes four ways: server clipping, beacon expressions, Ogg
   encapsulation and Magrat encapsulation, but others are possible.

8.1.1.  Server clipping

   Conceptually, the simplest solution to stream capture is just for the
   server to always start sending data from a dispatch point.  If the
   server streaming data can be aware of the internal structure of the
   BULK stream, it can make new clients wait until the next dispatch
   point to start sending them data.

8.1.2.  Beacon expressions

   To signal some of the dispatch points, the abstract yield of the BULK
   stream contains an expression whose byte pattern is unique in the
   stream, repeated frequently enough to minimize the length of bytes
   that a processing application must go through before achieving stream
   capture.  After reading that byte pattern, the processing application
   can start parsing BULK expressions.

Thierry                   Expires 31 March 2027                [Page 46]
Internet-Draft                    BULK1                   September 2026

   The problem with the idea of a beacon expression is that because BULK
   arrays can contain arbitrary bytes, no beacon expression exists that
   cannot appear in a BULK stream.  There are two solutions.  First, in
   a variety of situations, an application could know or preclude a byte
   pattern to appear in the stream, and chose that expression as beacon.
   Second, an application could choose a byte pattern that can only
   appear in a BULK array and, whenever a BULK array contains that byte
   pattern, represent it in the broadcasted BULK stream as the
   concatenation of its split across the beacon pattern.

   For example, if the beacon expression is 0xC3AABBCC, the expression
   #[8] 0x0000-C3AABBCC-0000 would become ( concat #[4] 0x0000-C3AA #[4]
   0xBBCC-0000 ).

8.1.3.  Ogg encapsulation

   The Ogg format[RFC3533] already provides an efficient mechanism for
   stream capture with a relatively low overhead.  It can stream and
   multiplex data from multiple media types.

   The BULK stream MUST begin with a version form, and the beginning of
   that form constitutes the codec identifier: ( version 1 (see
   Section 10.3).

   Ogg packets provided to the Ogg encoder MUST start and end at
   dispatch points.  This ensures than any complete Ogg packet provided
   to the processing application by the Ogg decoder is a valid BULK
   stream and can be parsed into zero or more expressions.  Granule
   position SHOULD be the number of expressions parsed in the abstract
   yield after parsing the Ogg packet.

8.1.4.  Magrat encapsulation

   The Magrat encapsulation is inspired by the Ogg format and has
   similar properties, except that it is less concerned with audio and
   video, is less powerful, and takes less space (hence the name).

   A Magrat broadcast stream is a sequence of Magrat forms.  The BULK
   stream that is encapsulated inside a Magrat broadcast stream is
   called the Magrat embedded stream.  This specification documents
   basic Magrat encapsulation.  In this version, a Magrat form has the
   following shape:

   ( values ( bulk {chunk}:Bytes ) {checksum}:Expr )

   This means that a processing application can look for the sequence of
   six bytes 0x01-1013-01-1009 to mark the beginning of a Magrat form.
   In case those bytes were in a BULK array, there are several features

Thierry                   Expires 31 March 2027                [Page 47]
Internet-Draft                    BULK1                   September 2026

   that work to verify that they actually started a Magrat form: first,
   they must be followed by an array {chunk}, a form end, a single
   expression {checksum} and another form end, second, {checksum} MUST
   be a digest that matches the bytes contained in {chunk}. Those bytes
   are called the Magrat chunk and they MUST be a valid BULK stream.  A
   protocol using Magrat encapsulation MAY specify which kind of
   checksum can be used.

   In the basic Magrat encapsulation, the concatenation of chunks is the
   embedded stream.  An extension of basic Magrat encapsulation MAY add
   metadata forms inside the chunk that are removed before the chunks
   are concatenated as the embedded stream.

8.2.  Full-duplex communication

8.2.1.  Separate channels

   At the BULK level, the simplest way to do full-duplex communication
   is when the underlying protocol can create separate channels for each
   direction.  Each agent streams its own BULK stream that is fully
   independent from the other agent's stream.  There is no restriction
   on the semantics used in the streams and an agent can use namespace
   associations and definitions to build a complex data model.

   Each agent is faced with the usual security considerations of BULK
   processing (see Section 9).  An agent faced with BULK data that
   crosses a safety threshold SHOULD stop processing the other agent's
   stream.  For that reason, a protocol using separate channels for BULK
   full-duplex communication SHOULD provide a way for an agent to end a
   channel.  This SHOULD include the reason for ending the channel, and
   if the agent offers the option to restart the channel from scratch.
   It MAY also include the option to restart the channel up to a
   previous safe dispatch point.

8.2.2.  Shared channel

   When two agents want to communicate over a bidirectional channel and
   reference data sent by each other, one naive way to do it would be to
   consider a virtual BULK stream acting as a kind of shared whiteboard.
   Each agent sending a BULK expression would add that expression to the
   white board.  When the agents have a mechanism to ensure some
   transactional safety, meaning that one agent cannot write without
   having properly read what the other agent had written, this is a
   valid option.  This can even work for more than two agents.

   For when this safety is not available, this specification defines
   BULK's basic full-duplex protocol.  In the basic protocol, the fact
   that it is used is an out-of-band information, which could be part of

Thierry                   Expires 31 March 2027                [Page 48]
Internet-Draft                    BULK1                   September 2026

   the underlying protocol statically or conveyed during connection
   establishment.  Another BULK full-duplex protocol could instead
   define a namespace to convey its use and some configuration
   parameters.

   In the basic protocol, the agent that initiates the communication is
   called the initiating agent and the other agent is called the
   responding agent.  The initiating agent and the responding agents
   each have a set of namespace markers designated for their use.  The
   initiating agent's designated markers are even numbers.  The
   responding agent's designated markers are odd numbers.  Agents MUST
   only associate immutable namespaces and MUST only associate them to
   their designated markers and MUST NOT associate a namespace to a
   marker that already is associated to a namespace.  Agents MUST only
   define names that don't already have a definition, and only to
   references whose namespace marker is in their designated markers.  It
   is a protocol error when any of those rules is broken by either
   agent.

   There is a separate virtual BULK stream associated with each agent.
   Each agent has a _pull position_, which is an offset in the other
   agent's virtual stream.  At the beginning of the exchange, each
   agent's pull position is 0, the beginning of the stream.

   Let Alice and Bob be two agents in a full-duplex communication.  When
   Alice streams a BULK expression A1 that only contains references with
   standard namespaces or Alice's designated markers, this expression A1
   is appended to Alice's virtual stream.  But when Alice streams a BULK
   expression A2 that contains one or several references with Bob's
   designated markers that got a definition in Bob's virtual stream,
   through namespace association or by definition, after Alice's pull
   position, those definitions are _pulled_, i.e. the content of Bob's
   virtual stream between Alice's pull position and the first dispatch
   point where all those references have a definition gets appended to
   Alice's virtual stream, Alice's pull position becomes that dispatch
   point, then A2 is appended to Alice's virtual stream.  When Alice
   streams a BULK expression A3 that contains one or several references
   with Bob's designated markers, but all those references got a
   definition before the pull position, only A3 is appended to Alice's
   virtual stream.

   If both agents have associated a single namespace to correct
   designated markers for each one, and both agents have independently
   defined a name that wasn't defined in the namespaces's immutable
   definition, it is a protocol error when either agent pulls the other
   agent's definition of that name.

Thierry                   Expires 31 March 2027                [Page 49]
Internet-Draft                    BULK1                   September 2026

   Whenever an agent sees a protocol error, it MUST end the connection.
   Before ending the connection, the agent MAY stream an expression
   shaped ( explain false {reason}:String ), in which case {reason} MUST
   contain a human-readable description of the issue that triggered the
   disconnect.

9.  Security Considerations

9.1.  Parsing

   Parsing a BULK stream is designed to be free of side-effects for the
   processing application, apart from storing the parsed results.

   Arrays in BULK carry their size, to avoid the need for escaping their
   content.  A malicious software, however, may announce an array with a
   size chosen to get an application to exhaust its available memory.
   When a BULK stream has been completely received, an array bigger than
   the remaining data is a parsing error.  When a BULK stream's size is
   not known in advance, the application SHOULD use a growable data
   structure.

   Evaluation opens up some known attacks that appear whenever a format
   provides a way to express abstraction, like the billion laughs
   attack.  As it is explained in Evaluation, an implementation MAY stop
   evaluation after a predefined number of evaluation steps.  As this
   has been demonstrated not to be sufficient to prevent attacks based
   on expansion, an implementation SHOULD also put predefined limits on
   the space that the concrete yield can take on disk or in memory.

   A processing application SHOULD use lazy immutable data structures to
   represent array concatenation and array indexing, as a defence
   against evaluaton attacks.  For example, in the billion laughs
   attack, the resulting concatenation would produce 9 lists of 10
   pointers and one actual array of 3 characters, instead of an array of
   3 billion characters.

   Applications MAY use out-of-band information to select size limits
   (like HTTP attributes), or a BULK namespace MAY provide hints.

9.2.  Forwarding

   When a processing application forwards all or part of the data in a
   BULK stream to another application, care must be taken if part of the
   forwarded data was not entirely recognized, as it could be used by an
   attacker to benefit from the authority the forwarding application has
   on the recipient of the data.

Thierry                   Expires 31 March 2027                [Page 50]
Internet-Draft                    BULK1                   September 2026

   If a protocol deems it necessary for applications to be able to
   forward data they don't fully understand, a known protection from
   that threat is the use of capability security, where the agent that
   provides the data to be forwarded must also provide the explicit
   authority that will be used after forwarding.  If the authority of
   the forwarding application is not used, it cannot be abused.

9.3.  Definitions

   The architecture of a processing application SHOULD ensure that a
   malicious agent cannot abuse authority given to it to define a
   namespace in order to modify associations in other namespaces.
   Depending on the use of data structures storing BULK expressions,
   this could amount to giving an attacker a way to manipulate the
   application's state.  See Appendix B for an example of architecture
   that is resistant to that kind of attack.

9.4.  Selectively parseable content

   It could be a security risk if a single BULK stream could be parsed
   into two different abstract yields by two conformant applications, so
   the evaluation of the whole stream cannot change whether some part
   that is designed to be selectively parseable is decoded or not.  For
   that reason, any side-effects in the selectively parseable
   expressions that affect how BULK expressions are evaluated (like
   namespace associations or definitions) MUST be isolated.

   For that security reason, there isn't a ( bulk-with-size Nat Expr )
   form to make the expression skippable, because it would open up that
   risk when the size given is not the actual size of the enclosed
   expression, accidentally or maliciously.

   Whenever BULK data is selectively parseable, it MUST be enclosed in a
   values form.

9.5.  BULK formats and protocols

   This specification doesn't address in too much detail the security
   considerations that a BULK format or protocol would need to include,
   because of the wide diversity of use cases for BULK.  But Appendix E
   gives some hints and resources on the subject.

10.  IANA Considerations

Thierry                   Expires 31 March 2027                [Page 51]
Internet-Draft                    BULK1                   September 2026

10.1.  Media type

   This specification defines two new media types, application/bulk and
   text/bulk.  Here are the informations for its registration to IANA
   [BCP13]:

10.1.1.  application/bulk

   Type name  application

   Subtype name  bulk

   Required parameters  N/A

   Optional parameters  N/A

   Encoding considerations  none, content is self-describing

   Security considerations  cf. Section 9

   Interoperability considerations  N/A

   Published specification  this document

   Applications that use this media type  the BARK manifest prototype

   Fragment identifier considerations  this specification defines no
      semantics for addressing the data with a fragment identifier; a
      future specification MAY define fragment identifier syntaxes to
      address the content by byte offset or the parsed results by their
      position in the abstract yield

   Additional information

      Magic numbers  the constraint to start any BULK file with a
         version form has the side-effect that classes of BULK streams
         can be identified by a sequence of bytes acting as "magic
         number", at offset 0:

         0x011000  any BULK stream

         0x01100081  a BULK stream of major version 1

         0x011000818002  a BULK stream of version 1.0

      File extensions  .bulk

      Structured type name suffix  [RFC6839]

Thierry                   Expires 31 March 2027                [Page 52]
Internet-Draft                    BULK1                   September 2026

         this specification defines a suffix +bulk for naming media
         types that use BULK as their core syntax

10.1.2.  text/bulk

   Type name  text

   Subtype name  bulk

   Required parameters  N/A

   Optional parameters  N/A

   Encoding considerations  content MUST be encoded in UTF-8 [STD63]

   Security considerations  cf. Section 9

   Interoperability considerations  N/A

   Published specification  this document, Appendix A

   Applications that use this media type  the BARK manifest prototype

   Fragment identifier considerations  this specification defines no
      semantics for addressing the data with a fragment identifier; a
      future specification MAY define fragment identifier syntaxes to
      address the content by byte offset or the parsed results by their
      position in the abstract yield

   Additional information

      Magic numbers  The text notation allows for arbitrary number and
         kinds of whitespaces around lexical elements, so there are no
         "magic numbers" as such but, in most cases, BULK streams in
         text notation will not have leading whitespace and use a single
         space within the first version form, so the first characters
         can identify classes of BULK streams:

         *  '( version' or '( bulk:version': any BULK stream

         *  '( version 1' or '( bulk:version 1': a BULK stream of major
            version 1

         *  '( version 1 0 )' or '( bulk:version 1 0 )': a BULK stream
            of version 1.0

      File extensions  .bulktext

Thierry                   Expires 31 March 2027                [Page 53]
Internet-Draft                    BULK1                   September 2026

10.2.  Link relation

   This specification defines a new link relation type, bulk-profile.
   Here are the informations for its registration to IANA [RFC8288]:

   Relation name  bulk-profile

   Description  This link target is the BULK stream that a processing
      application SHOULD use as the profile of the link's context.

   Reference  this document

10.3.  Ogg media mapping

   This specification defines a new Ogg logical bitstream type[RFC5334],
   for the media type application/bulk:

   Codec identifier  char[4]: '\x01\x10\x00\x81'

   Codecs parameter  bulk1

   For more details, see Section 8.1.3.

   A BULK-aware Ogg decoder could anticipate future BULK versions and
   recognize any version form conformant with Section 4.1 as codec
   identifier with the encompassing codecs parameter "bulk".

11.  Acknowledgements

   The original author of this specification read Erik Naggum's famous
   rant about XML (http://www.schnada.de/grapt/eriknaggum-xmlrant.html)
   several years before, and while forgotten as such for a time, it
   definitively was the seed that slowly bloomed into the design of
   BULK.  This format is dedicated to Erik.

   Unknowingly, work on BULK started just as CBOR[RFC8949] was in
   _Request for Last Call_ at IETF.  The early design goals of BULK and
   the design goals of CBOR had both significant differences and a large
   common ground.  It felt like an implicitly obvious choice to make the
   parser a simple state machine that could be implemented with a jump
   table but CBOR's inspiration was to make fast processing speed and
   low processing footprint explicit requirements.

   The idea to store together marking bits and a small argument in a
   marker byte was a direct inspiration from both CBOR and
   MessagePack[MsgPack] and it made BULK's syntax and implementation
   both drastically simpler.

Thierry                   Expires 31 March 2027                [Page 54]
Internet-Draft                    BULK1                   September 2026

12.  References

12.1.  Normative References

   [BCP13]    Best Current Practice 13,
              <https://www.rfc-editor.org/info/bcp13>.
              At the time of writing, this BCP comprises the following:

              Freed, N. and J. Klensin, "Multipurpose Internet Mail
              Extensions (MIME) Part Four: Registration Procedures",
              BCP 13, RFC 4289, DOI 10.17487/RFC4289, December 2005,
              <https://www.rfc-editor.org/info/rfc4289>.

              Freed, N., Klensin, J., and T. Hansen, "Media Type
              Specifications and Registration Procedures", BCP 13,
              RFC 6838, DOI 10.17487/RFC6838, January 2013,
              <https://www.rfc-editor.org/info/rfc6838>.

              Dürst, M.J., "Guidelines for the Definition of New Top-
              Level Media Types", BCP 13, RFC 9694,
              DOI 10.17487/RFC9694, March 2025,
              <https://www.rfc-editor.org/info/rfc9694>.

   [BCP14]    Best Current Practice 14,
              <https://www.rfc-editor.org/info/bcp14>.
              At the time of writing, this BCP comprises the following:

              Bradner, S., "Key words for use in RFCs to Indicate
              Requirement Levels", BCP 14, RFC 2119,
              DOI 10.17487/RFC2119, March 1997,
              <https://www.rfc-editor.org/info/rfc2119>.

              Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC
              2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174,
              May 2017, <https://www.rfc-editor.org/info/rfc8174>.

   [BCP18]    Best Current Practice 18,
              <https://www.rfc-editor.org/info/bcp18>.
              At the time of writing, this BCP comprises the following:

              Alvestrand, H., "IETF Policy on Character Sets and
              Languages", BCP 18, RFC 2277, DOI 10.17487/RFC2277,
              January 1998, <https://www.rfc-editor.org/info/rfc2277>.

   [IANA-Charsets]
              "IANA Charset Registry (archived at):",
              <http://www.iana.org/assignments/character-sets>.

Thierry                   Expires 31 March 2027                [Page 55]
Internet-Draft                    BULK1                   September 2026

   [RFC3533]  Pfeiffer, S., "The Ogg Encapsulation Format Version 0",
              RFC 3533, DOI 10.17487/RFC3533, May 2003,
              <https://www.rfc-editor.org/info/rfc3533>.

   [RFC5334]  Goncalves, I., Pfeiffer, S., and C. Montgomery, "Ogg Media
              Types", RFC 5334, DOI 10.17487/RFC5334, September 2008,
              <https://www.rfc-editor.org/info/rfc5334>.

   [RFC6839]  Hansen, T. and A. Melnikov, "Additional Media Type
              Structured Syntax Suffixes", RFC 6839,
              DOI 10.17487/RFC6839, January 2013,
              <https://www.rfc-editor.org/info/rfc6839>.

   [RFC8288]  Nottingham, M., "Web Linking", RFC 8288,
              DOI 10.17487/RFC8288, October 2017,
              <https://www.rfc-editor.org/info/rfc8288>.

   [STD63]    Internet Standard 63,
              <https://www.rfc-editor.org/info/std63>.
              At the time of writing, this STD comprises the following:

              Yergeau, F., "UTF-8, a transformation format of ISO
              10646", STD 63, RFC 3629, DOI 10.17487/RFC3629, November
              2003, <https://www.rfc-editor.org/info/rfc3629>.

12.2.  Informative references

   [Avro]     Cutting, D., "Apache Avro™ 1.7.4 Specification", February
              2013, <http://avro.apache.org/docs/1.7.4/spec.html>.

   [dhall-sec]
              Gonzalez, G., "Dhall Safety Guarantees",
              <https://docs.dhall-lang.org/discussions/Safety-
              guarantees.html>.

   [I-D.bormann-cbor-draft-numbers]
              Bormann, C., "Managing CBOR codepoints in Internet-
              Drafts", Work in Progress, Internet-Draft, draft-bormann-
              cbor-draft-numbers-08, 27 June 2026,
              <https://datatracker.ietf.org/doc/html/draft-bormann-cbor-
              draft-numbers-08>.

   [MsgPack]  Furuhashi, S., "MessagePack", <https://msgpack.org/>.

   [protobuf] "Protocol Buffers", July 2008,
              <https://developers.google.com/protocol-buffers/>.

Thierry                   Expires 31 March 2027                [Page 56]
Internet-Draft                    BULK1                   September 2026

   [RFC5234]  Crocker, D., Ed. and P. Overell, "Augmented BNF for Syntax
              Specifications: ABNF", STD 68, RFC 5234,
              DOI 10.17487/RFC5234, January 2008,
              <https://www.rfc-editor.org/info/rfc5234>.

   [RFC7540]  Belshe, M., Peon, R., and M. Thomson, Ed., "Hypertext
              Transfer Protocol Version 2 (HTTP/2)", RFC 7540,
              DOI 10.17487/RFC7540, May 2015,
              <https://www.rfc-editor.org/info/rfc7540>.

   [RFC8264]  Saint-Andre, P. and M. Blanchet, "PRECIS Framework:
              Preparation, Enforcement, and Comparison of
              Internationalized Strings in Application Protocols",
              RFC 8264, DOI 10.17487/RFC8264, October 2017,
              <https://www.rfc-editor.org/info/rfc8264>.

   [RFC8610]  Birkholz, H., Vigano, C., and C. Bormann, "Concise Data
              Definition Language (CDDL): A Notational Convention to
              Express Concise Binary Object Representation (CBOR) and
              JSON Data Structures", RFC 8610, DOI 10.17487/RFC8610,
              June 2019, <https://www.rfc-editor.org/info/rfc8610>.

   [RFC8949]  Bormann, C. and P. Hoffman, "Concise Binary Object
              Representation (CBOR)", STD 94, RFC 8949,
              DOI 10.17487/RFC8949, December 2020,
              <https://www.rfc-editor.org/info/rfc8949>.

   [RFC9839]  Bray, T. and P. Hoffman, "Unicode Character Repertoire
              Subsets", RFC 9839, DOI 10.17487/RFC9839, August 2025,
              <https://www.rfc-editor.org/info/rfc9839>.

   [Smile]    Saloranta, T., "Smile Data Format", September 2010,
              <https://github.com/FasterXML/smile-format-specification>.

   [Thrift]   Slee, M., Agarwal, A., and M. Kwiatkowski, "Thrift:
              Scalable Cross-Language Services Implementation", April
              2007, <http://thrift.apache.org/static/files/thrift-
              20070401.pdf>.

Appendix A.  Using the text notation as a format

   BULK's text notation can be used as a full-fledged format alongside
   BULK's binary syntax.

   The text format has a different trade-off.  On one hand, it is
   readily human-readable and it is easy to author in the absence of any
   tooling, because it is plain text.  On the other hand, it is less
   robust in several ways: it is more dependent on the processing

Thierry                   Expires 31 March 2027                [Page 57]
Internet-Draft                    BULK1                   September 2026

   application knowing the definitions of namespaces and packages, it is
   rigid with respect to encoding (it MUST be encoded in UTF-8), and it
   is limited and inefficient in its representation of arbitrary bytes
   and strings.  While parsing BULK text notation is a bit more involved
   than parsing binary BULK, the syntax of the text notation is still
   deliberately simple so as to limit even that part's processing
   footprint.

   Conceptually, parsing text notation involves translating the notation
   into binary BULK and then processing that.  Each lexical element of
   BULK's text notation is separated from other elements by whitespace.
   Apart from ([ and ]), delimiting an array containing a BULK stream,
   every lexeme can be immediately translated into its binary
   representation.

   Translating reference mnemonics involves knowing the mnemonics of
   previously imported namespaces, which means that as complete
   expressions are produced by the parser, they need to be evaluated.
   Reference mnemonics can appear without a namespace mnemonic if there
   aren't two imported namespaces that both use that mnemonic for a
   name.  When a reference mnemonic appears with a known namespace
   mnemonic but an unknown name mnemonic, a processing application MUST
   associate that mnemonic with the first name in that namespace that
   doesn't already have a mnemonic.  This makes it easy to author a
   namespace definition without having to manually number names.

Appendix B.  Robust namespace definition

   This constitutes a suggestion of architecture for a BULK processing
   application.  It has the advantage that an agent cannot modify the
   values of names to which it has not specifically been given
   authority.  This architecture doesn't ensure this property by
   checking the validity of definitions but by adhering to the Principle
   Of Least Authority, thus ensuring no false positives or TOCTOU race
   conditions.

   For each new context (including the abstract yield when parsing
   starts), the parser creates a new copy of each known namespace.
   These copies are available in this context to retrieve and define
   values.  It implements the lexical scoping of definitions on top of
   providing the robustness properties discussed here.

   By default, all namespaces created in a context are discarded at the
   end of this context.

   Of course, an implementation of the architecture presented here can
   be optimized compared to the abstract algorithm, for example by using
   copy-on-demand.

Thierry                   Expires 31 March 2027                [Page 58]
Internet-Draft                    BULK1                   September 2026

   Any namespace that is not a copy for its context but the object
   retained by the application afterwards, gives authority to make long-
   lasting definitions.  A namespace that is stored by the processing
   application after evaluating a BULK stream is called a lasting
   namespace.

   Note that there are two ways to define a namespace in a BULK stream:
   using only the definition in the ( define ( namespace {…} ) form, or
   using this (possibly empty) definition as modified by other
   definitions (in the same BULK stream or not).  The former is called
   the _initial definition_, the latter the _amended definition_.

B.1.  Complete authority

   When the amended definitions of all namespaces constitute lasting
   namespaces, it means that the evaluation of the BULK stream can
   modify any existing namespace.  This level of authority might be
   useful for a handful of privileged BULK streams acting as
   configuration of the application (e.g. to achieve reverse aliasing,
   see Appendix C).

B.2.  Selective authority

   A number of lasting namespaces are included for the abstract yield.
   Their unique identifiers are agreed out-of-band.  The disadvantage of
   this solution is that it needs prior agreement on the definable
   namespaces.  This may be a safer way than complete authority to
   achieve reverse aliasing (see Appendix C).

B.3.  Open authority

   Any namespace definition for a unique identifier unknown to the
   processing application triggers the creation of a lasting namespace.

   The disadvantage of this solution is that it opens a denial of
   service vulnerability.  If Bob is a processing application and Carol
   and Dave are agents communicating with Bob with an open authority,
   Dave can prevent Carol from defining a namespace if it manages to
   know the unique identifier and to start a communication with Bob
   before Carol.

   If an agent uses a secure way to create unique identifiers, this
   solution is both flexible and safe (the burden is not on the BULK
   processing application).  This specification thus encourages the use
   of open authority restricted to verifiable namespaces (in which case
   several agents can present the same definition to a processing
   application without conflict).

Thierry                   Expires 31 March 2027                [Page 59]
Internet-Draft                    BULK1                   September 2026

   A processing application could have a configuration setting to select
   what will generate a lasting namespace:

   unrestricted open authority  any time a namespace is defined with an
      identifier that was previously unknown, either its initial or
      amended definition constitutes a lasting namespace; this is the
      most lax open authority, most flexible but also most open to
      misuse and issues

   collision-free open authority  any time a namespace is defined with
      an identifier that was previously unknown and the identifier
      relies explicitly on an algorithm that the processing application
      deems giving a high enough guarantee that identifiers are unique,
      either the initial or amended definition of that namespace
      constitutes a lasting namespace

   immutable open authority  any time an immutable namespace is defined,
      its initial definition constitutes a lasting namespace

   It is RECOMMENDED to use immutable open authority by default, as
   several agents can safely present the same definition to a processing
   application without conflict.

Appendix C.  Forward compatibility

   BULK makes it possible to create new versions of vocabularies that
   encompass previous versions, in a way that minimizes implementation
   complexity.

   The first tool is aliasing: reuse names and values from existing
   namespaces, even in bootstrapping namespaces:

   ( define ( namespace ( newhash:shake128 {newhashid} ) 20 )
     ([ ( version 1 0 )
     ( ( namespace 20 ) )
     ( mnemonic ( namespace 20 ) "newhash" )
     ( explain ( namespace 20 ) "The new, shiny hash namespace!" )
     ( mnemonic newhash:shake128 "shake128" )
     ( import 21 ( namespace ( oldhash:shake128 {oldhashid} ) ) )
     ( define newhash:shake128 oldhash:shake128 ) ]) )

   With this, new namespaces can be created and applications don't need
   to change the existing code.

   One possible downside with aliasing is that if the number of aliasing
   namespaces grow, you might end up with the implementation of an
   important namespace scattered across a bunch of aliased legacy
   namespaces.  Also, the definition of the new namespace is tied to the

Thierry                   Expires 31 March 2027                [Page 60]
Internet-Draft                    BULK1                   September 2026

   old one, which means that you need to keep the old definition around
   for the lifetime of the new one.  To prevent those issues, a second
   tool is to reverse the direction of aliasing: all the implementation
   lives in the current namespace, cohesively, and its definition can be
   used on its own, and the old namespace is aliased to the new:

   ( import 20 ( namespace ( oldhash:shake128 {oldhashid} ) ) )
   ( import 21 ( namespace ( newhash:shake128 {newhashid} ) ) )
   ( define oldhash:shake128 newhash:shake128 )

   Following the Principle of Least Authority, it should not be possible
   by default for the evaluation of any BULK stream to make lasting
   modifications to existing namespaces.

   One obvious design would be for the application to have a privileged
   storage for reverse aliasing namespace definitions, with those
   namespaces given complete authority or, better yet, each being given
   selective authority for a specific existing namespace (see
   Appendix B).  Where this could still not be deemed safe enough,
   reverse aliasing of namespaces could be defined in the application's
   code.

Appendix D.  Arity-carrying forms

   Sometimes a vocabulary will include forms that can contain an
   arbitrary number of expressions.  When such a form is used in postfix
   bytecode, the simplest solution is just to use a nested postfix form:

   ( define ( arity ) ( 2 black white comment ) )
   ( postfix
     game
     1 2 black
     ( postfix alternative
       "white tried an unorthodox opening" 3 4 white comment
       "a more classical opening would be" 8 9 white comment )
     2 3 black
     ( postfix alternative
       "white played a bad move" 4 5 white comment
       "white could have played a decent move" 5 6 white comment
       "white could have played a great move" 5 7 white comment ) )

   The nested postfix form costs 4 bytes, compared to an equivalent
   postfix bytecode.

Thierry                   Expires 31 March 2027                [Page 61]
Internet-Draft                    BULK1                   September 2026

   If those 4 bytes add up to too much space through repetition, an
   application could define a form for the sole purpose of assigning it
   an arity, while the evaluation of the arity-carrying form would just
   replace it with the original one.  For example, after evaluating the
   postfix bytecode transformation and the resulting form of the last
   expression of

   ( define alt/2 alternative )
   ( define alt/3 alternative )
   ( define ( arity ) ( 2 black white comment alt/2 ) ( 3 alt/3 ) )
   ( postfix
     game
     1 2 black
     "white tried an unorthodox opening" 3 4 white comment
     "a more classical opening would be" 8 9 white comment
     alt/2
     2 3 black
     "white played a bad move" 4 5 white comment
     "white could have played a decent move" 5 6 white comment
     "white could have played a great move" 5 7 white comment
     alt/3
     )

   it would be transformed into

  ( game
    ( black 1 2 )
    ( alternative
      ( comment "white tried an unorthodox opening" ( white 3 4 ) )
      ( comment "a more classical opening would be" ( white 8 9 ) ) )
    ( black 2 3 )
    ( alternative
      ( comment "white played a bad move" ( white 4 5 ) )
      ( comment "white could have played a decent move" ( white 5 6 ) )
      ( comment "white could have played a great move" ( white 5 7 ) ) )
    ( white 4 5 ) )

   Such an arity-carrying form costs 10 or 13 bytes to be usable when it
   is added to an existing form defining arities.  Which means that
   compared to the nested postfix form, it pays for itself if it is used
   only 3 or 4 times.

Appendix E.  The difference between BULK and BULK formats

   BULK aims at being a useful framework for a wide variety of formats,
   including low-level communication protocols, higher-level RPC or REST
   APIs, media files, rich documents, archives and efficient
   serialization of existing data models (like XML, JSON or RDF).

Thierry                   Expires 31 March 2027                [Page 62]
Internet-Draft                    BULK1                   September 2026

   This had several implications on its design.

E.1.  Purposefully open: for generality

   As such, BULK imposes no constraints on what kind of data can be
   represented.  One is free to design a BULK format that mandates the
   use of EBCDIC.  In keeping with [BCP18], BULK chooses UTF-8 as the
   default string encoding and its core namespace only permits
   designating encodings from the IANA charset registry[IANA-Charsets].
   But a BULK vocabulary would be free to define a new windows-codepage
   form to use Windows Codepages instead.

   Any protocol designed to transport human-readable text should be
   aware of [RFC9839], but BULK could be used to transport text emitted
   in a terminal, making full use of control characters, or even to
   create a file with examples of ill-formed UTF-8 strings containing
   surrogates.  As such, BULK doesn't limit what code points are allowed
   in UTF-8 strings, or any other Unicode encoding.

   BULK doesn't include a grammar to define BULK formats, but a BULK
   grammar vocabulary, that could encode ABNF[RFC5234] or CDDL[RFC8610]
   in BULK, would do well to add the ability to express "Unicode
   Scalars", "XML Characters" and "Unicode Assignables" from [RFC9839],
   as well as classes and profiles from PRECIS[RFC8264].

E.2.  Purposefully limited: for safety

   For the same reason, BULK parsing and evaluation needed to be secure
   by default and feature a security model that would be safe enough
   that it can be a secure foundation almost everywhere.

   This is why BULK syntax and the BULK core namespace can't directly
   express notions like the inclusion of an outside BULK stream, or
   referencing a file or URI to access.  This is also why this
   specification limits the discoverability of bootstrapping namespaces
   and strongly limits discoverabilty of non immutable namespaces and
   packages.

   A BULK format that can express such a dangerous combination of
   actions as reading files and making network connections SHOULD
   carefully consider the attack surface they present and the threat
   models for the format's use cases, and explain those in detail in the
   format's documentation.  Dhall's Safety Guarantees[dhall-sec] are a
   prime example.

   The BULK core namespace can't directly express Turing complete
   functions, and not even functions that can take different execution
   paths depending on their arguments.  The functions that can be

Thierry                   Expires 31 March 2027                [Page 63]
Internet-Draft                    BULK1                   September 2026

   expressed can only use their arguments in a static transformation, by
   design.  While this drastically limit BULK's expressivity, it also
   drastically limit the attack surface on what arbitrary input can make
   the BULK parser or evaluator do, while still offering a decent power
   of abstraction (e.g. one could write a substitution function to
   unpack an array containing an IPv6 header into its fields, but not a
   substitution function to unpack an IPv6 extension header according to
   its type).

   The goal is that a BULK protocol or format designer should be able to
   trust that if they use BULK in accordance with the safety
   recommendations of this specification, they don't need to carefully
   weigh the benefits of evaluation vs. its dangers, like it has been
   the case with a couple of previous formats.

Appendix F.  Marking and extending media types with BULK

   One possible use of BULK is the ability to add metadata around a
   file.  The most basic metadata is the media type and its parameters.
   Although many media types can be expressed with a file extension,
   this usually doesn't encode parameters like the charset used for a
   plain text file format.  This is why most spreadsheet software
   present the user with a preview of a few rows while asking them for
   the encoding, when importing CSV data.

   A media type vocabulary would make it possible to encode parameters
   and provide names for some known parameter values.

   While a .md file extension only encodes the media type text/markdown,
   a short BULK header could encode the full media type text/markdown;
   charset=ISO-8859-15; variant=GFM:

   ( version 1 0 )
   ( import 20 ( package ( shake128 #[8] 0xDABBED01 ) 2 ) )
   ( markdown ( iana-charset 111 ) "GFM" )

   While a .csv file extension only encodes the media type text/csv, a
   short BULK header could encode the full media type text/csv;
   charset=UTF-8; header=absent

   ( version 1 0 )
   ( import 20 ( package ( shake128 #[8] 0xDABBED01 ) 2 ) )
   ( csv ( iana-charset 106 ) csv-header-absent )

   Media type metadata could be mixed with other metadata:

Thierry                   Expires 31 March 2027                [Page 64]
Internet-Draft                    BULK1                   September 2026

   ( version 1 0 )
   ( import 20 ( package ( shake128 #[8] 0xDABBED02 ) 4 ) )
   ( description
     ( media-type ( csv ( iana-charset 106 ) csv-header-present ) )
     ( licence cc0 ) )

Author's Address

   Pierre Thierry
   Comonad Dev
   Email: pierre@comonad.dev

Thierry                   Expires 31 March 2027                [Page 65]