Data Vault

👤 vitorhugoze 📦 v1.0.18 ⭐ 4.3 ⬇️ 708 下载
📊 数据分析 免费

📖 技能介绍


name: data-vault version: 1.0.18 description: "Persist and retrieve structured data using the Lance columnar format. Use when you need to store, query, or analyze data across sessions — such as saving skill outputs, tracking conversation context, storing research data, or building knowledge bases. After installing the requirements it's ready to use. Triggers on: 'store this data', 'save to persistant storage', 'persist information', 'remember this', 'store for later', 'query my data', 'analyze stored data', 'persist data'." author: Vitor Hugo Zeferino metadata: openclaw: requires: bins: - python3 # declare uv and pip as required binaries - uv - pip install: # Bootstrap pip if missing - kind: "shell" cmd: "python3 -m ensurepip --upgrade || true" label: "Ensure pip is installed"

        # Bootstrap uv if missing
        - kind: "shell"
          cmd: "pip install --upgrade uv || true"
          label: "Install uv if missing"

        # Install pylance
        - kind: "uv"
          type: "pip"
          package: "pylance"
          label: "Install pylance (Lance columnar format) via uv"

        # Install pandas
        - kind: "uv"
          type: "pip"
          package: "pandas"
          label: "Install pandas via uv"

Data Vault

Installation

uv pip install pylance pandas

A persistent data store using the Lance columnar format for fast ML data access.

Quick Start

# List all datasets and their metadata
python3 scripts/command.py list-datasets-info

# Create a dataset
python3 scripts/command.py create-dataset <name> <field1> <field2> ...

# Append data
python3 scripts/command.py append-to-dataset <name> <value1> <value2> ...

# Read all records from a dataset
python3 scripts/command.py read-dataset <name>

Note: list-datasets-info shows dataset metadata (schema, field types, record count) — it does not return the actual data rows. Use read-dataset to retrieve records.

Storage Location

DataSets are created and stored on the current path '.'

Critical Behavior: Data Type Strictness

⚠️ Lance is strict about data types — they CANNOT change after the first record

When you append the first record to a dataset, Lance infers the data type for each field. All subsequent records MUST use the same types.

Example — this FAILS:

# First record: age as STRING
append-to-dataset users "John" "25" "john@test.com"

# Second record: age as INTEGER (will FAIL!)
append-to-dataset users "Jane" 30 "jane@test.com"
# Error: `age` should have type large_string but type was int64

Correct approach — maintain consistent types:

# First record: age as STRING
append-to-dataset users "John" "25" "john@test.com"

# Second record: age as STRING
append-to-dataset users "Jane" "30" "jane@test.com"

Why This Matters

Unlike traditional databases that may coerce types, Lance rejects type mismatches. If you store numbers as strings initially, you must always pass strings. Plan your schema carefully.

Initialization Workflow

When starting a session, always initialize by listing existing datasets first:

# This command returns ALL datasets with their structure
python3 scripts/command.py list-datasets-info

Example output:

{
    "skill": "data-vault",
    "operation": "list_datasets_info",
    "status": "success",
    "data": [
        {
            "dataset_name": "users",
            "path": "/data/users",
            "fields": ["name", "age", "email"],
            "field_types": {
                "_id": "large_string",
                "_updated_at": "timestamp[us]",
                "name": "large_string",
                "age": "large_string",
                "email": "large_string"
            },
            "record_count": 2,
            "columns": ["id", "_updated_at", "name", "age", "email"],
            "last_updated": "2026-03-21T17:57:44.595628"
        }
    ],
    "error": null
}

Understanding field_types

State Meaning
{} (empty) Dataset exists but no records yet — types not yet defined
populated Types are locked — appends must match

Important: If field_types is empty, the first append will define types. Be deliberate about the first record's types.

Commands Reference

Create Dataset

python3 scripts/command.py create-dataset <name> <field1> <field2> ...

Creates a metadata entry. Fields have no types until first append.

Append Record

python3 scripts/command.py append-to-dataset <name> <value1> <value2> ...

Appends one record. Types are inferred from first record.

Batch Append

python3 scripts/command.py batch-append-to-dataset <name> '<json-array>'

Example: batch-append-to-dataset users '[["Alice", "22", "alice@test.com"], ["Bob", "35", "bob@test.com"]]'

Update Record

python3 scripts/command.py update-dataset-record <name> <record_id> <value1> <value2> ...

Updates fields for a specific record by ID.

Delete Record

python3 scripts/command.py delete-dataset-record <name> <record_id>

List All Datasets

python3 scripts/command.py list-datasets

Get Dataset Info

python3 scripts/command.py get-dataset-info <name>

Returns schema, field types (if data exists), and record count.

List All Datasets with Full Info

python3 scripts/command.py list-datasets-info

Recommended for initialization. Returns all datasets with complete metadata.

Get Dataset Path

python3 scripts/command.py get-dataset-path-info <name>

Backup Dataset

python3 scripts/command.py backup-dataset <name> <backup_path>

Count Records

python3 scripts/command.py count-records <name>

Read All Records

Returns all records from the dataset as a list of objects.

python3 scripts/command.py read-dataset <name>

Drop Dataset

Requires confirmation if have not created a backup beforehand.

Delete the entire dataset and its metadata.

python3 scripts/command.py drop-dataset <name>

Internal fields available in every dataset:

发现更多技能插件,请访问7w4.net。

Field Type Description
_id string UUID — unique record identifier
_updated_at timestamp When the record was last inserted or updated

List Records (Paginated)

python3 scripts/command.py list-records <name> --limit 10 --offset 0

Returns records with optional pagination.

Get Single Record

python3 scripts/command.py get-record <name> <record_id>

Retrieves a specific record by its UUID.

Get Dataset Info

python3 scripts/command.py get-dataset-info <name>

Returns schema, field types (if data exists), and record count.

Response Format

All commands return JSON:

{
  "skill": "data-vault",
  "operation": "<operation_name>",
  "status": "success|error",
  "data": <result_data_or_null>,
  "error": <error_message_or_null>
}

Internal Fields

Every dataset automatically includes:

  • _id — UUID for each record
  • _updated_at — timestamp of last insert/update

These are managed automatically — when appending, only provide your defined fields.

Data Type Inference

Lance infers types from the first record:

Python Type Lance Type
"string" large_string
25 (int) int64
25.5 (float) float64
True/False bool

CLI caveat: When passing via command line, all values are strings. To ensure integer types, initialize with actual integers in a script rather than CLI.

Tips

  1. Initialize at session start: Run list-datasets-info to understand what data already exists
  2. Plan your schema: First record determines types for the entire dataset
  3. Use batch append when adding multiple records: More efficient than individual appends

Requirements

Dependencies are declared in frontmatter (metadata.openclaw.install) and handled by the OpenClaw install system via uv. The Python packages required are:

  • pylance — The Lance columnar format library.

    ⚠️ Naming note: Despite the PyPI package being named pylance, the library is imported as import lance in Python code. This is the official Lance project naming convention — it is NOT the VS Code "pylance" language server. See lance.org for details.

  • pandas — Data manipulation

🤖 AI 评测

这个 Skill 质量不错,文档写得非常详细清楚,连数据类型不能乱改这种容易踩坑的地方都特意提醒了。功能挺齐全的,存数据、查数据、删数据、备份都能做。但有编码问题导致说明文件显示乱码,而且依赖包的安装和导入名字不一样,新手可能搞不清楚。总体来说是款实用工具,主要缺点就是文档有小瑕疵、新手入门有点门槛。

📊 多维度评分

适应性4.5
规范性4.2
有效性4.4
可靠性4.4
可信度4.3

📁 包含文件 (9 个)

📄 README.md 2.8 KB
📄 SKILL.md 8.9 KB
📄 _meta.json 130 B
📄 requirements.txt 296 B
📄 scripts/__init__.py 0 B
📄 scripts/command.py 5.6 KB
📄 scripts/manage.py 7.7 KB
📄 scripts/read.py 3.9 KB
📄 scripts/write.py 5.2 KB