lance-format
Deep reference for Lance v13 columnar format, its Rust crates, and pylance - file encodings, table format, indexes, schema evolution, time travel. Use when building on the Lance crates or reading .lance datasets, not the LanceDB product.
Rank
62
Safety
84
Downloads
2.6k
Updated
Oct 9, 2026
Version
0.20.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.6K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.6K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.20.0release · observed Sep 17, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17bp3v1hm1dnkzey0c9tfh02183j0y5:lance-format- Install using `clawhub skill install s17bp3v1hm1dnkzey0c9tfh02183j0y5:lance-format` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/tenequm/lance-format before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/snapshot"
Documentation
CLAWHUB
160,000 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: lance-format description: Deep reference for Lance v13 columnar format, its Rust crates, and pylance - file encodings, table format, indexes, schema evolution, time travel. Use when building on the Lance crates or reading .lance datasets, not the LanceDB product. metadata: version: "0.20.0" categories: "development, integrations" topics: "lance, columnar-format, vector-search, rust, lakehouse" upstream: "lance-format/[email protected]" openclaw: homepage: https://github.com/tenequm/skills/tree/main/skills/lance-format emoji: "🗄️" --- # Lance v13 reference Lance is an open columnar format for multimodal AI - "a columnar data format that is 100x faster than Parquet for random access." It is not one format but a stack of interoperating specs: a **file format**, a **table format**, **index formats**, **catalog specs**, and a **namespace client spec**. The Rust workspace at `lance-format/lance` implements all of them plus Python (`pylance`) and Java bindings. This skill tracks **`v13.0.0-beta.4`** (the `lance-format/lance` git tag), the current development frontier; **`v12.0.0`** is the stable pin, released 2026-09-17. Pin against tags, not `main` - Lance ships beta tags every few days and `next`-format encodings can change. Version landscape below. Three layers of reference, load what the task needs: - **The deep reference** - any concrete schema, parameter, proto, or constraint. Split by topic: | File in `references/` | Covers | Sections | |------|--------|----------| | `format-file.md` | What Lance is, the 26 crates, file format, data types | 1-4 | | `format-table.md` | Dataset layout, manifests, fragments, schema evolution, versioning/tags/branches, row IDs, transactions + OCC, MemWAL | 5-10 | | `indexes.md` | Vector / scalar / FTS / geo indexes, distributed builds | 11-12 | | `ops.md` | Object store, capability matrix, source map | 13, 15, 16 | | `changelog-v7-v13.md` | The full v7 -> v13 delta | 14 | Cross-references written as "section N" resolve through `references/lance-reference.md`. - `references/performance.md` - ALL performance guidance. Part A routes to the official text and adds the source-derived changes upstream has not documented; Part B is field-verified remote-storage practice. Load for any performance, tuning, maintenance-cost, or "why is this slow" question. - `references/docs/` - a **verbatim mirror of the official docs** (`docs/src` at the tracked tag): every guide, quickstart, and format spec, unedited. Load when you need the full official text. Directory map below. `references/maintenance.md` covers refreshing this skill against a new upstream tag. ## Lance vs LanceDB These are two different things and conflating them produces wrong answers. - **Lance** - the format and engine. The `lance-format/lance` repo; the `lance` /`lance-*` Rust crates; `pylance`. It gives you datasets, the file/table format, indexes, commits, scans. Consumed directly by DuckDB, P
_meta.json
{
"ownerId": "kn76gpsgjw5chv0xvzbzcb8cxn81x46r",
"slug": "lance-format",
"version": "0.20.0",
"publishedAt": 1789677222469
}references/changelog-v7-v13.md
# Lance changelog - v7 -> v12 (section 14) Part of the Lance v13 reference (`lance-format/[email protected]`). Citations are `path:line` relative to the repo root; build a permalink as `https://github.com/lance-format/lance/blob/v13.0.0-beta.4/<path>`. Line numbers drift between tags - treat them as approximate. Cross-references written as "section N" use the original 16-section numbering; `lance-reference.md` maps every number to its file. **Release-line shape.** The major is bumped by a bot, not a human: `ci/publish_beta.sh:65,87` re-roots at `MAJOR+1` whenever any PR since the release root carries the GitHub `breaking-change` label (`ci/check_breaking_changes.py:31`). The marker is the **label**, not a conventional-commit `!` - of the 13 labeled PRs in the v11 window (#8024, #8025, #8026, #8027, #8028, #8051, #8095, #8159, #8172, #8188, #8206, #8347, #8360) only two carry `!` in the subject. It has now fired on two consecutive lines: `9.1.0-beta.*` -> `10.0.0-beta.*` (2026-07-23), then `10.1.0-beta.*` -> `11.0.0-beta.*` (2026-08-05, `649076df1 chore: bump to 11.0.0-beta.1 based on breaking change detection`). So **neither `v9.1.0` nor `v10.1.0` was ever released**, and `v10.1.0-beta.2` is the direct ancestor of `v11.0.0-beta.1`, one bump commit apart. The re-root renumbers in place: `release-root/10.1.0-beta.N` and `release-root/11.0.0-beta.N` point at the same commit (`ee0a60d0c`), both recording `Base: 10.0.0-rc.1`. **`v10.0.0` final was tagged on 2026-08-08** - an annotated, PGP-signed tag ("Release version 10.0.0") on `release/v10.0`, one commit past `v10.0.0-rc.3` (2026-08-02). That branch forked at `v10.0.0-beta.7` and took one substantive backport (`10d0c9f2e fix: backport encoding and FTS fixes to release/v10.0`, #8146). It is **not** an ancestor of `main` - finals are cut on `release/vX.Y` branches, so that is normal. `v10.0.0-beta.7` **is** an ancestor of `v11.0.0-beta.16`, but `v10.0.0-rc.3` and `v10.0.0` are not. **`v10.0.0` is the stable pin** (2026-08-08, superseding `v9.0.1`), and it is what GitHub Releases marks `Latest`. `v9.0.1` (2026-08-06, superseding `v9.0.0`, 2026-07-24) shipped with five sibling patch finals that day - `v8.0.1`, `v7.1.0`, `v6.1.0`, `v4.0.2`, `v3.0.2` - each on its own `release/vX.Y` branch. `v5.0.0` still has no final despite `v5.0.0-rc.2`. crates.io publishes **finals only** (`max_stable_version` = `10.0.0`; no 11.x, and the only pre-release among ~186 versions is the ancient `0.0.1-alpha0`); PyPI `pylance` is likewise at `10.0.0`. So any beta pin is a git dependency; beta artifacts publish to fury.io (`.github/workflows/publish-beta.yml:114`) under the renamed org, `https://pypi.fury.io/lance-format`. ## Contents - [The v7.1.0-beta.1 delta](#the-v710-beta1-delta) - [The v7.1.0-beta.2 delta](#the-v710-beta2-delta) - [The v7.1.0-beta.2 -> v7.2.0-beta.5 delta](#the-v710-beta2---v720-beta5-delta) - [The v7.2.0-beta.5 -> v8.0.0-beta.9 delta (major-version boundary)](#the-v720-beta5---v800-beta9-del
references/docs/format/file/encoding.md
# Lance Encoding Strategy The encoding strategy determines how array data is encoded into a disk page. The encoding strategy tends to evolve more quickly than the file format itself. ## Older Encoding Strategies The 0.1 and 2.0 encoding strategies are no longer documented. They were significantly different from future encoding strategies and describing them in detail would be a distraction. ## Terminology An array is a sequence of values. An array has a data type which describes the semantic interpretation of the values. A layout is a way to encode an array into a set of buffers and child arrays. A buffer is a contiguous sequence of bytes. An encoding describes how the semantic interpretation of data is mapped to the layout. An encoder converts data from one layout to another. Data types and layouts are orthogonal concepts. An integer array might be encoded into two completely different layouts which represent the same data.  ### Data Types Lance uses a subset of Arrow's type system for data types. An Arrow data type is both a data type and an encoding. When writing data Lance will often normalize Arrow data types. For example, a string array and a large string array might end up traveling down the same path (variable width data). In fact, most types fall into two general paths. One for fixed-width data and one for variable-width data (where we recognize both 32-bit and 64-bit offsets). At read time, the Arrow data type is used to determine the target encoding. For example, a string array and large string array might both be stored in the same layout but, at read time, we will use the Arrow data type to determine the size of the offsets returned to the user. There is no requirement the output Arrow type matches the input Arrow type. For example, it is acceptable to write an array as "large string" and then read it back as "string". ## Search Cache The search cache is a key component of the Lance file reader. Random access requires that we locate the physical location of the data in the file. To do so we need to know information such as the encoding used for a column, the location of the page, and potentially other information. This information is collectively known as the "search cache" and is implemented as a basic LRU cache. We define a "initialization phase" which is when we load the various indexing information into the search cache. The cost of initialization is assumed to be amortized over the lifetime of the reader. When performing full scans (i.e. not random access), we should be able to ignore the search cache and sometimes can avoid loading it entirely. We _do_ want to optimize for cold scans as the initialization phase is often not amortized over the lifetime of the reader. ## Structural Encoding The first step in encoding an array is to determine the structural encoding of the array. A structural encoding breaks the data into smaller units which can be independentl
references/docs/format/file/index.md
# Lance File Format
The Lance file format is a columnar container optimized for cloud object stores, random access, and Arrow-native processing. It deliberately focuses on page layout and encoding mechanics, while leaving table semantics and search structures to higher layers.
## Design Goals
### No Row Groups
Lance does not use Parquet-style row groups. Each column may have its own number of pages, which keeps column data in large storage-friendly chunks regardless of schema width and avoids coupling scanner partitioning to physical file layout.
### Random-Access-Friendly Encoding
Pages are designed so readers can fetch contiguous row ranges with a small and predictable number of I/O operations. This is important for selective filters, point lookups, vector-search follow-up reads, and ML training workloads that sample rows non-sequentially.
### Functional Decomposition
The file layer does not bundle table-level statistics or query-side indices into the base file structure. Those capabilities are defined as separate index formats so they can evolve independently of the core file container.
## File Structure
A Lance file is a container for tabular data. The data is stored in "disk pages". Each disk page contains some rows
for a single column. There may be one or more disk pages per column. Different columns may have different numbers of
disk pages. Metadata at the end of the file describes where the pages are located and how the data is encoded.

!!! Note
This page describes the container specification. We also have a set of default encodings that are used to encode
data into disk pages. See the [Encoding Strategy](encoding.md) page for more details.
### Disk Pages
Disk pages are designed to be large enough to justify a dedicated I/O operation, even on cloud storage, typically several megabytes. Using a larger page size may reduce the number of I/O operations required to read a file, but it also increases the amount of memory required to write the file. In practice, very large page sizes are not useful when high speed reads are required because large contiguous reads need to be broken into smaller reads for performance (particularly on cloud storage). As a result, a default of 8MB is recommended for the page size and should yield ideal performance on all storage systems.
Disk pages should not generally be opaque. It is possible to read a portion of a disk page when a subset of the rows are
required. However, the specifics of this process depend on the column encoding which is described in a later section.
### No Row Groups
Unlike similar formats, there is no "row group" concept, only pages. We believe the concept of row groups to be
fundamentally harmful to performance. If the row group size is too small then columns will be split into "runt pages" which yield poor read performance on cloud storage. If the row group size is too large then a file writer will needactivepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/tenequm/skills/lance-format",
"sourceUrl": "https://clawhub.ai/tenequm/skills/lance-format",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T12:51:59.743Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T12:51:59.743Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.6K downloads",
"href": "https://clawhub.ai/tenequm/lance-format",
"sourceUrl": "https://clawhub.ai/tenequm/lance-format",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T12:51:59.743Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.20.0",
"href": "https://clawhub.ai/tenequm/lance-format",
"sourceUrl": "https://clawhub.ai/tenequm/lance-format",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-17T20:33:42.469Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.20.0",
"description": "Updated lance-format from 0.19.0 to 0.20.0. Changes: - modified `CHANGELOG.md` - modified `SKILL.md` - renamed `references/changelog-v7-v12.md` -> `references/changelog-v7-v13.md` - modified `references/docs/format/file/versioning.md` - modified `references/docs/format/index/index.md` - modified `references/docs/format/index/system/frag_reuse.md` - modified `references/docs/format/index/vector/index.md` - modified `references/docs/format/table/row_id_lineage.md` - modified `references/docs/format/table/transaction.md` - modified `references/docs/format/table/versioning.md` - modified `references/docs/guide/object_store.md` - modified `references/docs/guide/performance.md`",
"href": "https://clawhub.ai/tenequm/lance-format",
"sourceUrl": "https://clawhub.ai/tenequm/lance-format",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-17T20:33:42.469Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
