Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
50 changes: 50 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,56 @@ trainer = Trainer(model=model, train_dataset=dataset["train"], eval_dataset=data
trainer.train()
```

### Image-Text-to-Text Datasets

Some AgML datasets (e.g. `Project-AgML/AgroOmni`) are too large to fit the standard `load_dataset()`
archive handling, which fully extracts every archive to disk before reading it. These datasets instead store images
as raw relative paths inside zip shards, alongside a single shared `path -> shard` index, so no extraction step is
needed and no per-archive metadata has to be duplicated. Because of that, they aren't `load_dataset()`-compatible,
and are loaded with `loadImageTextToTextDataset` instead:

```python
from agml import loadImageTextToTextDataset

# Whole dataset, split-aware:
ds = loadImageTextToTextDataset(
"Project-AgML/AgroOmni",
token=HF_TOKEN
)

ds["train"][0] # {'id', 'messages', 'raw_metadata', 'images': [PIL.Image, ...]}

# Single split:
train_ds = loadImageTextToTextDataset(
"Project-AgML/AgroOmni",
split="train",
token=HF_TOKEN
)

# Multi-config repo (each config under its own folder): pass `config` to select one.
ds = loadImageTextToTextDataset(
"Project-AgML/MIRAGE",
config="MMST_Standard",
token=HF_TOKEN,
)
```

`loadImageTextToTextDataset(repo_id, config=None, split=None, cache_dir=None, token=None, cache_max_bytes=...)` returns a `DatasetDict`
when `split` is omitted, or a single `Dataset` when a split name is given. Splits are auto-discovered from the
metadata parquet filenames (e.g. `train-0000-of-0001.parquet` becomes `train`), the same convention `load_dataset()`
itself uses; if no split can be inferred, every row is placed under a single `train` split. For repos with multiple
configs, each living under its own folder, `config` is required, if it's omitted the error message lists the
configs available in that repo.

Image bytes are only read the first time a row is actually accessed (`ds["train"][0]`, a slice, or a batch), never
eagerly over the whole dataset, and repeated access to the same image is cached. The `ImageTextToTextShardStore`
backing that lazy decoding is attached to the returned dataset as `ds.store`, in case you want to inspect its cache,
call `ds.store.getBytes(path)` directly, or free its open shard file handles and cached bytes (`ds.store.close()`).
The cache is bounded by total size, 256 MiB by default; pass `cache_max_bytes=` to change it, or `0` to disable it.

If you're using this with a PyTorch `DataLoader`, prefer `num_workers > 0`. Each worker process opens its
own shard handles (and keeps its own byte cache), whether workers are forked or spawned.

## Public Datasets

AgML contains a wide variety of public datasets from various locations across the world:
Expand Down
4 changes: 3 additions & 1 deletion agml/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,5 +38,7 @@ def _setup():
del _setup # noqa


# There are no top-level imported functions or classes, only the modules.
# Besides the modules, `loadImageTextToTextDataset` is exposed at the top level
# since it is the entry point for loading AgML image-text-to-text datasets from the Hub.
from . import backend, data, io, synthetic, viz
from .data import loadImageTextToTextDataset
1 change: 1 addition & 0 deletions agml/data/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,3 +20,4 @@
from .tools import coco_to_bboxes, convert_bbox_format
from .hf_loader import HuggingFaceDataLoader
from .multispectral_hf_loader import MultispectralDataLoader
from .image_text_to_text_loader import ImageTextToTextShardStore, loadImageTextToTextDataset
Loading
Loading