Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Icarus-Dataset

A unified multi-modal curriculum dataset for evolutionary neural architecture search (ENAS), published as one Hugging Face dataset plus a tiny vendorable loader. This repository is the generator: it sources each rung, normalizes everything to one structural schema, carves support/query, and writes a self-contained, streamable dataset directory.

What it is

Every row of the dataset is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. A Field carries its tensor plus a descriptor (axes, value_type, n_classes, value_range, mask) that an encoder/loss reads instead of inspecting raw shape, so the dataset stays structural and the encoder is swappable. The rungs form an 18-step difficulty ladder across modalities:

Rung Name Task
1 XOR 2-bit XOR, the smallest non-linearly-separable problem
2 Parity-N parity of a padded, mask-delimited bit vector
3 Two-spirals which of two interleaved spirals a 2-D point lies on
4 Pole (Markov) next-state prediction of a cart-pole from a short window
5 Double-pole (no velocity) two poles, velocities hidden (a memory benchmark)
6 MNIST / Fashion-MNIST 10-way grayscale image classification
7 CIFAR-10 / CIFAR-100 natural color image classification
8 NB360 ecg single-lead ECG rhythm classification
9 NB360 satellite land-cover from a satellite time series
10 NB360 ninapro hand gesture from surface-EMG
11 NB360 spherical spherically-projected image classification
12 NB360 cosmic per-pixel cosmic-ray segmentation
13 NB360 darcy_flow Darcy-flow PDE solution-operator regression
14 NB360 psicov protein residue-distance regression (variable length)
15 NB360 fsd50k multi-label audio tagging from a mel-spectrogram
16 NB360 deepsea multi-label chromatin marks from a DNA window
17 RAVEN / PGM abstract visual analogy (pick the completing panel)
18 ARC-AGI v1 few-shot grid program induction (native train/test split)

Explorer

image image image image

Layout

icarus/     vendorable runtime (torch + datasets ONLY): schema (types), codec, batching, loader
etl/        build pipeline: per-rung adapters, parquet writer, dataset card, vendor bundler, CLI
explore/    Rich terminal viewer
examples/   reference Level0Encoder + loss_fn (swappable; powers the end-to-end smoke test)
tests/      schema round-trip, loader semantics, per-rung ETL, end-to-end training, vendor-bundle checks

The runtime in icarus/ is the only thing meant to be vendored downstream, and it depends on torch + datasets alone. The build (etl/) flattens that runtime (plus the reference encoder) into a single self-contained icarus.py written into the dataset directory at build time, with a PEP 723 header so a consumer can uv run it without installing this repo. Keeping the loader and encoder in one module is deliberate: Python enums compare by identity, so a split loader/encoder would need the two files side by side to compose; one file removes that coupling.

Setup

Python 3.12 only, managed with uv. No notebooks.

uv sync

Build

uv run build --rungs 1,3,6 --out .dataset_build   # a few rungs to a local directory
uv run build --all --dry-run                       # count + round-trip check, write nothing
uv run build --all --out /Volumes/Pickles/Icarus_dataset   # full build (rung 17 / PGM is large)

A non-dry-run build writes, per rung, rung_<N>/shard-*.parquet, plus README.md (the Hub dataset card with a configs: block), SOURCES.md, LICENSE.md, .gitattributes, and the vendored icarus.py. The build is idempotent and seeded, so re-runs reproduce identical output.

Explore

uv run explore                     # reads /Volumes/Pickles/Icarus_dataset by default
ICARUS_DATA=.dataset_build uv run explore

An arrow-key terminal viewer that renders each rung's tasks (images, grids, time series, vectors) with a faithful tensor preview and a per-rung description, reading the local parquet directly.

Use the published dataset

See the dataset card on the Hub for full examples. In short, stream rows and reconstruct tasks with the shipped loader, or use the vendored IcarusDataset for whole Task objects:

from datasets import load_dataset
from icarus import IcarusDataset, deserialize_task  # the single icarus.py shipped with the dataset

row = next(iter(load_dataset("Ardea/Icarus_dataset", name="rung_6", streaming=True, split="train")))
task = deserialize_task(row)

dataset = IcarusDataset(rungs=(3, 6, 18), n_tasks=100, n_samples=50, hf_repo="Ardea/Icarus_dataset")
task = dataset[0]

Publishing the dataset

Publishing is manual via git + LFS (the build emits a .gitattributes that routes *.parquet through LFS, so the big files are handled for you):

# 1. Create an EMPTY dataset repo at huggingface.co/new-dataset (Ardea/Icarus_dataset; do NOT add a README)
# 2. Full build into the output directory
uv run build --all --out /Volumes/Pickles/Icarus_dataset

# 3. Publish from the output directory
cd /Volumes/Pickles/Icarus_dataset
git init
git lfs install
git add -A
git commit -m "Icarus dataset release"
git remote add origin https://huggingface.co/datasets/Ardea/Icarus_dataset
git push -u origin main   # auth: use a Hugging Face access token as the password

Rung 17 (PGM) is very large; if git-LFS struggles at that scale, hf upload-large-folder is the Hugging Face recommended alternative for the same directory.

Development

uv run pytest            # tests (single test: uv run pytest path::test_name)
uv run ruff check        # lint
uv run ruff format       # format (line length 180)
uv run ty check          # type checking (astral ty, not mypy)

License

MIT with Attribution, Copyright (c) 2026 Ardea AI Corp. See LICENSE.md. The published dataset redistributes third-party sources (rungs 6-18), each under its own upstream license recorded in SOURCES.md and the dataset card; comply with the upstream license of any rung you use.

About

Progressive ladder dataset

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages