Claim this agent
agentCLAWHUBUnverified

lance-format

Deep reference for Lance v13 columnar format, its Rust crates, and pylance - file encodings, table format, indexes, schema evolution, time travel. Use when building on the Lance crates or reading .lance datasets, not the LanceDB product.

OpenClaw

Rank

62

Safety

84

Downloads

2.6k

Updated

Oct 9, 2026

Version

0.20.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.6K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.6K downloadsadoption · observed Oct 9, 2026
Latest release
0.20.0release · observed Sep 17, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17bp3v1hm1dnkzey0c9tfh02183j0y5:lance-format
  1. Install using `clawhub skill install s17bp3v1hm1dnkzey0c9tfh02183j0y5:lance-format` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/tenequm/lance-format before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/snapshot"

Documentation

CLAWHUB

160,000 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: lance-format
description: Deep reference for Lance v13 columnar format, its Rust crates, and pylance - file encodings, table format, indexes, schema evolution, time travel. Use when building on the Lance crates or reading .lance datasets, not the LanceDB product.
metadata:
  version: "0.20.0"
  categories: "development, integrations"
  topics: "lance, columnar-format, vector-search, rust, lakehouse"
  upstream: "lance-format/[email protected]"
  openclaw:
    homepage: https://github.com/tenequm/skills/tree/main/skills/lance-format
    emoji: "🗄️"
---

# Lance v13 reference

Lance is an open columnar format for multimodal AI - "a columnar data format that is 100x
faster than Parquet for random access." It is not one format but a stack of interoperating
specs: a **file format**, a **table format**, **index formats**, **catalog specs**, and a
**namespace client spec**. The Rust workspace at `lance-format/lance` implements all of them
plus Python (`pylance`) and Java bindings.

This skill tracks **`v13.0.0-beta.4`** (the `lance-format/lance` git tag), the current
development frontier; **`v12.0.0`** is the stable pin, released 2026-09-17. Pin against tags, not
`main` - Lance ships beta tags every few days and `next`-format encodings can change. Version
landscape below.

Three layers of reference, load what the task needs:

- **The deep reference** - any concrete schema, parameter, proto, or constraint. Split by topic:

  | File in `references/` | Covers | Sections |
  |------|--------|----------|
  | `format-file.md` | What Lance is, the 26 crates, file format, data types | 1-4 |
  | `format-table.md` | Dataset layout, manifests, fragments, schema evolution, versioning/tags/branches, row IDs, transactions + OCC, MemWAL | 5-10 |
  | `indexes.md` | Vector / scalar / FTS / geo indexes, distributed builds | 11-12 |
  | `ops.md` | Object store, capability matrix, source map | 13, 15, 16 |
  | `changelog-v7-v13.md` | The full v7 -> v13 delta | 14 |

  Cross-references written as "section N" resolve through `references/lance-reference.md`.
- `references/performance.md` - ALL performance guidance. Part A routes to the official text and
  adds the source-derived changes upstream has not documented; Part B is field-verified
  remote-storage practice. Load for any performance, tuning, maintenance-cost, or "why is this
  slow" question.
- `references/docs/` - a **verbatim mirror of the official docs** (`docs/src` at the tracked
  tag): every guide, quickstart, and format spec, unedited. Load when you need the full official
  text. Directory map below.

`references/maintenance.md` covers refreshing this skill against a new upstream tag.

## Lance vs LanceDB

These are two different things and conflating them produces wrong answers.

- **Lance** - the format and engine. The `lance-format/lance` repo; the `lance` /`lance-*`
  Rust crates; `pylance`. It gives you datasets, the file/table format, indexes, commits,
  scans. Consumed directly by DuckDB, P

_meta.json

{
  "ownerId": "kn76gpsgjw5chv0xvzbzcb8cxn81x46r",
  "slug": "lance-format",
  "version": "0.20.0",
  "publishedAt": 1789677222469
}

references/changelog-v7-v13.md

# Lance changelog - v7 -> v12 (section 14)

Part of the Lance v13 reference (`lance-format/[email protected]`). Citations are `path:line`
relative to the repo root; build a permalink as
`https://github.com/lance-format/lance/blob/v13.0.0-beta.4/<path>`. Line numbers drift between
tags - treat them as approximate. Cross-references written as "section N" use the original
16-section numbering; `lance-reference.md` maps every number to its file.

**Release-line shape.** The major is bumped by a bot, not a human: `ci/publish_beta.sh:65,87`
re-roots at `MAJOR+1` whenever any PR since the release root carries the GitHub
`breaking-change` label (`ci/check_breaking_changes.py:31`). The marker is the **label**, not a
conventional-commit `!` - of the 13 labeled PRs in the v11 window (#8024, #8025, #8026, #8027,
#8028, #8051, #8095, #8159, #8172, #8188, #8206, #8347, #8360) only two carry `!` in the
subject. It has now fired on two
consecutive lines: `9.1.0-beta.*` -> `10.0.0-beta.*` (2026-07-23), then `10.1.0-beta.*` ->
`11.0.0-beta.*` (2026-08-05, `649076df1 chore: bump to 11.0.0-beta.1 based on breaking change
detection`). So **neither `v9.1.0` nor `v10.1.0` was ever released**, and `v10.1.0-beta.2` is
the direct ancestor of `v11.0.0-beta.1`, one bump commit apart. The re-root renumbers in place:
`release-root/10.1.0-beta.N` and `release-root/11.0.0-beta.N` point at the same commit
(`ee0a60d0c`), both recording `Base: 10.0.0-rc.1`.

**`v10.0.0` final was tagged on 2026-08-08** - an annotated, PGP-signed tag ("Release version
10.0.0") on `release/v10.0`, one commit past `v10.0.0-rc.3` (2026-08-02). That branch forked at
`v10.0.0-beta.7` and took one substantive backport (`10d0c9f2e fix: backport encoding and FTS
fixes to release/v10.0`, #8146). It is **not** an ancestor of `main` - finals are cut on
`release/vX.Y` branches, so that is normal. `v10.0.0-beta.7` **is** an ancestor of
`v11.0.0-beta.16`, but `v10.0.0-rc.3` and `v10.0.0` are not.

**`v10.0.0` is the stable pin** (2026-08-08, superseding `v9.0.1`), and it is what GitHub
Releases marks `Latest`. `v9.0.1` (2026-08-06, superseding `v9.0.0`, 2026-07-24) shipped with
five sibling patch finals that day - `v8.0.1`, `v7.1.0`, `v6.1.0`, `v4.0.2`, `v3.0.2` - each on
its own `release/vX.Y` branch. `v5.0.0` still has no final despite `v5.0.0-rc.2`. crates.io
publishes **finals only** (`max_stable_version` = `10.0.0`; no 11.x, and the only pre-release
among ~186 versions is the ancient `0.0.1-alpha0`); PyPI `pylance` is likewise at `10.0.0`. So
any beta pin is a git dependency; beta artifacts publish to fury.io
(`.github/workflows/publish-beta.yml:114`) under the renamed org,
`https://pypi.fury.io/lance-format`.

## Contents

- [The v7.1.0-beta.1 delta](#the-v710-beta1-delta)
- [The v7.1.0-beta.2 delta](#the-v710-beta2-delta)
- [The v7.1.0-beta.2 -> v7.2.0-beta.5 delta](#the-v710-beta2---v720-beta5-delta)
- [The v7.2.0-beta.5 -> v8.0.0-beta.9 delta (major-version boundary)](#the-v720-beta5---v800-beta9-del

references/docs/format/file/encoding.md

# Lance Encoding Strategy

The encoding strategy determines how array data is encoded into a disk page. The encoding strategy tends to evolve
more quickly than the file format itself.

## Older Encoding Strategies

The 0.1 and 2.0 encoding strategies are no longer documented. They were significantly different from future encoding
strategies and describing them in detail would be a distraction.

## Terminology

An array is a sequence of values. An array has a data type which describes the semantic interpretation of the values.
A layout is a way to encode an array into a set of buffers and child arrays. A buffer is a contiguous sequence of
bytes. An encoding describes how the semantic interpretation of data is mapped to the layout. An encoder converts
data from one layout to another.

Data types and layouts are orthogonal concepts. An integer array might be encoded into two completely different
layouts which represent the same data.

![Multiple Encodings](../../images/encoding_v_array.png)

### Data Types

Lance uses a subset of Arrow's type system for data types. An Arrow data type is both a data type and an encoding.
When writing data Lance will often normalize Arrow data types. For example, a string array and a large string array
might end up traveling down the same path (variable width data). In fact, most types fall into two general paths. One
for fixed-width data and one for variable-width data (where we recognize both 32-bit and 64-bit offsets).

At read time, the Arrow data type is used to determine the target encoding. For example, a string array and large
string array might both be stored in the same layout but, at read time, we will use the Arrow data type to determine
the size of the offsets returned to the user. There is no requirement the output Arrow type matches the input Arrow
type. For example, it is acceptable to write an array as "large string" and then read it back as "string".

## Search Cache

The search cache is a key component of the Lance file reader. Random access requires that we locate the physical
location of the data in the file. To do so we need to know information such as the encoding used for a column,
the location of the page, and potentially other information. This information is collectively known as the "search
cache" and is implemented as a basic LRU cache. We define a "initialization phase" which is when we load the various indexing information into the search cache. The cost of initialization is assumed to be amortized over the lifetime
of the reader.

When performing full scans (i.e. not random access), we should be able to ignore the search cache and sometimes
can avoid loading it entirely. We _do_ want to optimize for cold scans as the initialization phase is often not
amortized over the lifetime of the reader.

## Structural Encoding

The first step in encoding an array is to determine the structural encoding of the array. A structural encoding
breaks the data into smaller units which can be independentl

references/docs/format/file/index.md

# Lance File Format

The Lance file format is a columnar container optimized for cloud object stores, random access, and Arrow-native processing. It deliberately focuses on page layout and encoding mechanics, while leaving table semantics and search structures to higher layers.

## Design Goals

### No Row Groups

Lance does not use Parquet-style row groups. Each column may have its own number of pages, which keeps column data in large storage-friendly chunks regardless of schema width and avoids coupling scanner partitioning to physical file layout.

### Random-Access-Friendly Encoding

Pages are designed so readers can fetch contiguous row ranges with a small and predictable number of I/O operations. This is important for selective filters, point lookups, vector-search follow-up reads, and ML training workloads that sample rows non-sequentially.

### Functional Decomposition

The file layer does not bundle table-level statistics or query-side indices into the base file structure. Those capabilities are defined as separate index formats so they can evolve independently of the core file container.

## File Structure

A Lance file is a container for tabular data. The data is stored in "disk pages". Each disk page contains some rows
for a single column. There may be one or more disk pages per column. Different columns may have different numbers of
disk pages. Metadata at the end of the file describes where the pages are located and how the data is encoded.

![Format Overview](../../images/file_high_level_overview.png)

!!! Note

    This page describes the container specification. We also have a set of default encodings that are used to encode
    data into disk pages. See the [Encoding Strategy](encoding.md) page for more details.

### Disk Pages

Disk pages are designed to be large enough to justify a dedicated I/O operation, even on cloud storage, typically several megabytes. Using a larger page size may reduce the number of I/O operations required to read a file, but it also increases the amount of memory required to write the file. In practice, very large page sizes are not useful when high speed reads are required because large contiguous reads need to be broken into smaller reads for performance (particularly on cloud storage). As a result, a default of 8MB is recommended for the page size and should yield ideal performance on all storage systems.

Disk pages should not generally be opaque. It is possible to read a portion of a disk page when a subset of the rows are
required. However, the specifics of this process depend on the column encoding which is described in a later section.

### No Row Groups

Unlike similar formats, there is no "row group" concept, only pages. We believe the concept of row groups to be
fundamentally harmful to performance. If the row group size is too small then columns will be split into "runt pages" which yield poor read performance on cloud storage. If the row group size is too large then a file writer will need
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/tenequm/skills/lance-format",
      "sourceUrl": "https://clawhub.ai/tenequm/skills/lance-format",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T12:51:59.743Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T12:51:59.743Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.6K downloads",
      "href": "https://clawhub.ai/tenequm/lance-format",
      "sourceUrl": "https://clawhub.ai/tenequm/lance-format",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T12:51:59.743Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.20.0",
      "href": "https://clawhub.ai/tenequm/lance-format",
      "sourceUrl": "https://clawhub.ai/tenequm/lance-format",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-17T20:33:42.469Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tenequm-lance-format/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.20.0",
      "description": "Updated lance-format from 0.19.0 to 0.20.0. Changes: - modified `CHANGELOG.md` - modified `SKILL.md` - renamed `references/changelog-v7-v12.md` -> `references/changelog-v7-v13.md` - modified `references/docs/format/file/versioning.md` - modified `references/docs/format/index/index.md` - modified `references/docs/format/index/system/frag_reuse.md` - modified `references/docs/format/index/vector/index.md` - modified `references/docs/format/table/row_id_lineage.md` - modified `references/docs/format/table/transaction.md` - modified `references/docs/format/table/versioning.md` - modified `references/docs/guide/object_store.md` - modified `references/docs/guide/performance.md`",
      "href": "https://clawhub.ai/tenequm/lance-format",
      "sourceUrl": "https://clawhub.ai/tenequm/lance-format",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-17T20:33:42.469Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to lance-format and adjacent AI workflows.