Agent Skills: TileDB-VCF

Efficient storage and retrieval of genomic variant data using TileDB. Scalable VCF/BCF ingestion, incremental sample addition, compressed storage, parallel queries, and export capabilities for population genomics.

UncategorizedID: K-Dense-AI/claude-scientific-skills/tiledbvcf

Install this agent skill to your local

pnpm dlx add-skill https://github.com/K-Dense-AI/scientific-agent-skills/tree/HEAD/skills/tiledbvcf

Skill Files

Browse the full folder contents for tiledbvcf.

Download Skill

Loading file tree…

skills/tiledbvcf/SKILL.md

Skill Metadata

Name
tiledbvcf
Description
Stores and retrieves genomic variant calls with TileDB-VCF. Use for indexed single-sample VCF/BCF ingestion, incremental cohorts, region and sample queries, streaming results, allele statistics, QC, and VCF/BCF export locally or through TileDB Cloud.

TileDB-VCF

When to use

Use for cohort variant storage, incremental sample ingestion, interval queries, and exporting subsets for downstream genomics. TileDB-VCF stores records; it does not perform joint variant calling, association testing, normalization, or population structure adjustment. Dataset size alone does not determine whether cloud execution is appropriate: benchmark the intended sample, region, and attribute workload.

This skill targets released TileDB-VCF 0.40.3. Native Python/CLI checks used two synthetic single-sample VCFs on macOS ARM64. Cloud client 0.14.4 contracts were checked against its released wheel and mocked dispatch; hosted queries and jobs were not executed. See installation and verification notes.

Install and verify the native stack

The tiledbvcf Python distribution is not published on PyPI at this review. The official tiledb Conda channel supplies tiledbvcf-py 0.40.3 for macOS ARM64, macOS x86-64, and Linux x86-64, with Python 3.9-3.12 builds. tiledb on PyPI is TileDB-Py and does not install TileDB-VCF. Native Apple Silicon no longer requires forcing CONDA_SUBDIR=osx-64.

For the tested macOS ARM64 environment, pins below avoid versioned-library import failures in the otherwise successful current Conda solve:

conda create -n tiledb-vcf -c tiledb -c conda-forge \
  python=3.12 tiledbvcf-py=0.40.3 \
  azure-core-cpp=1.16.2 azure-storage-blobs-cpp=12.16.0 \
  azure-storage-files-datalake-cpp=12.14.0 capnproto=1.4.0 c-blosc2=2.23.1
conda activate tiledb-vcf
python -c 'import tiledbvcf; print(tiledbvcf.version)'
tiledbvcf version

The equivalent Micromamba solve/install was executed. These extra ABI pins are a verified macOS ARM64 workaround, not a claim about all platforms. Preserve the resolved environment for production. Other-platform installation and official Docker images are alternatives documented upstream, not tested here.

Prepare and create a cohort

  1. Record reference assembly, contig naming/lengths, callers, normalization rules, sample identity, and file checksums. Do not mix assemblies or assume 1 and chr1 are equivalent. Confirm compatible headers across inputs.
  2. Each input must contain one sample, be coordinate sorted, and have an index. Use BGZF-compressed VCF or BCF with a matching .csi/.tbi. Inspect bcftools query -l sample1.vcf.gz; filenames are not sample identifiers.
  3. Create the dataset explicitly, then ingest. Opening mode="w" alone does not create its schema. Materialize frequently read fields at creation time.
  4. Validate a small known interval and a round-trip export before scaling.

For existing sorted plain-text VCFs (run once per input; output names must be new):

bcftools view -Oz -o sample1.vcf.gz sample1.vcf
bcftools index -c sample1.vcf.gz
bcftools view -Oz -o sample2.vcf.gz sample2.vcf
bcftools index -c sample2.vcf.gz

The following local examples were exercised with sample IDs S1 and S2, contig chr1, and coordinates 1-100; substitute the verified IDs and intervals in real data.

import tiledbvcf

uri = "cohort"
config = {"sm.compute_concurrency_level": "2", "sm.io_concurrency_level": "2"}
with tiledbvcf.Dataset(uri, mode="w", tiledb_config=config) as ds:
    ds.create_dataset(extra_attrs=["fmt_GT", "fmt_DP"])
    ds.ingest_samples(
        ["sample1.vcf.gz"], threads=2, total_memory_budget_mb=512
    )

# Incremental ingestion uses the existing dataset; do not call create_dataset again.
with tiledbvcf.Dataset(uri, mode="w", tiledb_config=config) as ds:
    ds.ingest_samples(
        ["sample2.vcf.gz"], threads=2, total_memory_budget_mb=512
    )

The schema-v4 data array has contig, start-coordinate, and sample dimensions; headers and optional statistics are separate arrays in the dataset group. Fields not explicitly materialized remain in the INFO/FORMAT payloads. Query names such as pos_start need not match raw storage column names.

Upstream supports parallel thread/process ingestion. Assign distinct sample work and coordinate lifecycle/maintenance operations; do not blindly re-ingest samples. resume=True supports interrupted ingestion, not arbitrary duplicate correction.

Query completely and interpret coordinates correctly

| Surface | Coordinate convention | | --- | --- | | Python regions=["chr1:10-14"] | 1-based, closed interval | | Returned pos_start, pos_end | 1-based, inclusive record endpoints | | BED input and query_bed_start, query_bed_end | 0-based, half-open | | read_variant_stats() / read_allele_count() column pos in 0.40.3 | 0-based; add one before joining to VCF POS |

Queries return overlapping records, not just records starting inside the interval. A deletion spanning positions 10-14 appears in chr1:14-14, with pos_start=10. For BED [13,14), the equivalent string is chr1:14-14. The Python region parser requires explicit contig:start-end; bare "chr1" is rejected in this release. Use a known contig length for a whole-contig query.

cfg = tiledbvcf.ReadConfig(memory_budget_mb=128, tiledb_config=config)
attrs = ["sample_name", "contig", "pos_start", "pos_end", "alleles", "fmt_GT"]
with tiledbvcf.Dataset(uri, cfg=cfg) as ds:
    assert {"S1", "S2"}.issubset(ds.samples())
    assert set(attrs).issubset(ds.attributes())
    for batch in ds.read_iter(
        attrs=attrs, regions=["chr1:1-100"], samples=["S1", "S2"]
    ):
        print(batch)  # Replace with a bounded consumer or partitioned output.
  • read() returns a pandas DataFrame; read_arrow() returns a PyArrow Table. Either can return only the first memory-limited batch. Continue with continue_read() / continue_read_arrow() until read_completed(), or use read_iter() for ordinary DataFrame queries. Avoid concatenating all batches when the result cannot fit in memory.
  • ReadConfig(memory_budget_mb=...) controls reads; use ingest_samples(total_memory_budget_mb=...) for writes. A read budget is not a process RSS cap. limit truncates total output; it is not a pagination size.
  • sample_partition=(0, 2) selects partition zero of two. Likewise, region_partition=(index, number_of_partitions) partitions query regions; these tuples are not coordinate bounds, tile extents, or sample capacities.
  • 0.40.3 retains BED selection state when a later call omits bed_file. Open a fresh Dataset for an independent query when changing BED/region inputs.
  • fmt_GT contains allele indices: 0 refers to REF, positive values index ALT, and -1 means missing. Preserve multiallelic identity and ploidy; do not treat every positive integer as a biallelic dosage or missing calls as reference. The integer list alone does not encode the original phased GT string.

Export and validate

from pathlib import Path

Path("exported").mkdir(exist_ok=True)
with tiledbvcf.Dataset(uri, cfg=cfg) as ds:
    ds.export(
        samples=["S1"], regions=["chr1:1-100"],
        output_format="v", output_dir="exported",
    )

Python export formats: v VCF, z compressed VCF, u uncompressed BCF, b compressed BCF. merge=False writes per-sample files; merge=True requires a combined output_path. A combined export is not joint calling. Re-index exported files before indexed downstream access. Compare sample IDs, contigs, record counts, alleles, INFO/FORMAT values, and missing/phased genotypes; semantic round-trip preservation does not imply byte-identical compression or headers.

The CLI accepts positional sample paths or --samples-file (one URI per line), not a comma-separated --samples argument:

tiledbvcf create --uri cli_cohort
tiledbvcf store --uri cli_cohort --threads 2 --total-memory-budget-mb 512 \
  -- sample1.vcf.gz sample2.vcf.gz
tiledbvcf list --uri cli_cohort
tiledbvcf stat --uri cli_cohort
tiledbvcf export --uri cli_cohort --regions chr1:1-100 \
  --sample-names S1,S2 -Ot --tsv-fields 'SAMPLE,CHR,POS,REF,ALT,F:GT' \
  --output-path variants.tsv

-- protects positional paths from variable-length options such as --tiledb-config. Check expected artifacts and sample counts as well as the exit code: an invalid store command returned exit zero while printing an error in the reviewed CLI. TSV fields use F:GT for FORMAT/GT and I:DP for INFO/DP.

Cohort statistics and QC

with tiledbvcf.Dataset(uri, cfg=cfg) as ds:
    stats = ds.read_variant_stats(regions=["chr1:1-100"], drop_ref=True)
    stats["vcf_pos"] = stats["pos"] + 1
    print(stats[["contig", "vcf_pos", "alleles", "ac", "an", "af"]])

qc = tiledbvcf.sample_qc(uri, samples=["S1"], config=config)
print(qc)

Statistics arrays are enabled by default at creation; older/disabled arrays may not support these operations. read_variant_stats has no samples argument. The older read_allele_frequency(dataset_uri, region) wrapper accepts a single region and delegates to the deprecated singular argument; prefer the Dataset method above. sample_qc uses dataset_uri, not uri, as its first argument.

The ac, an, and af values summarize the ingested cohort, not a selected ancestry or query sample subset. A read(samples=[...], set_af_filter=">0.6") filter still used cohort AF in the tested release. Compute subgroup frequencies from correctly selected calls with explicit missingness, ploidy, callable-region, and gVCF reference-block policies. Use an appropriate called-allele denominator, not universally 2 * number_of_samples; a missing VCF record is not proof of a homozygous-reference call. QC metrics and sparse storage do not establish GWAS readiness or scientific validity.

Object storage and TileDB Cloud

Direct storage uses s3://bucket/path, azure://container/path, or gcs://bucket/path, supported by the installed TileDB backend and its provider credentials. This does not automatically distribute the computation. Use the provider's credential chain or scoped configuration, and verify access to both the dataset and source indexes. A TileDB Cloud token is distinct from bucket credentials. Remote storage examples below are illustrative, source-verified, and not authenticated end-to-end tests.

Install tiledb-cloud==0.14.4 in a compatible environment. Its life-sciences extra adds TileDB-SOMA, not TileDB-VCF. Supply TILEDB_REST_TOKEN through the execution environment before importing the client; TILEDB_REST_HOST selects a custom deployment if required. Do not embed tokens in code or print configuration.

# Illustrative: requires an accessible registered cohort and a billing namespace.
import tiledb.cloud
import tiledb.cloud.vcf

cloud_cfg = tiledb.cloud.Config()
with tiledbvcf.Dataset("tiledb://my-namespace/cohort", tiledb_config=cloud_cfg) as ds:
    sample_names = ds.samples()

result = tiledb.cloud.vcf.read(
    dataset_uri="tiledb://my-namespace/cohort",
    config=cloud_cfg, attrs=["sample_name", "pos_start", "fmt_GT"],
    regions=["chr1:1-100"], samples=sample_names[:2],
    num_region_partitions=1, namespace="my-namespace", max_workers=2,
)
frame = result.to_pandas()  # read returns an Arrow table, not a DataFrame.

Distributed reads assemble an Arrow result and may use significant worker/client memory. Bound the requested regions/samples and worker count; partitioning does not make the final concatenated result memory-free. This is SDK task execution, not a paginated VCF REST-list API.

# Illustrative, mutating cloud job: use only for an authorized ingestion task.
submission = tiledb.cloud.vcf.ingest(
    dataset_uri="s3://my-bucket/cohort",
    sample_list_uri="s3://my-bucket/inputs/sample-uris.txt",
    namespace="my-namespace", acn="registered-storage-role",
    register_name="cohort", max_samples=2,
    ingest_resources={"cpu": "2", "memory": "4Gi"},
)
print(submission["graph_id"])

Use exactly one of search_uri, sample_list_uri, or metadata_uri; file-search patterns apply to search_uri. Registration requires the access credential name (acn). vcf.ingest submits asynchronously and returns {"status": "started", "graph_id": ...}; track that graph to terminal success and verify ingested samples before claiming completion. There is no released ingest_vcf_dataset(source=..., output=...) API. Pricing, availability, security controls, and deployment obligations require the current account/service terms.

Official references