Research · Dataset release

OevuAI Image Training Corpus v1.0.0

A small, carefully built image–text dataset for studying training-data construction: every pair is documented, every release is checksummed, and the method behind it is written down.

79image–text pairs
298 MBtotal size
57 + 22dense-tag / structured bilingual pairs
SHA-256published per release
Corpus sample 00 Corpus sample 01 Corpus sample 04 Corpus sample 06 Corpus sample 08 Corpus sample 11 Corpus sample 03 Corpus sample 07
Samples from the corpus. The full 79-pair set and its annotation files ship in the release.

What's inside

  • DataImages with per-pair text annotations
  • StatsCategory counts, tag histogram, resolution distributions
  • ProvenanceSource and construction method documented
  • IntegritySHA-256 checksums for verification

Release record & integrity

Version1.0.0 — 2026-10-06
Archive297,379,369 bytes (298 MB)
SHA-25642694f4ea3f722a2…
AnnotationHuman-annotated, single annotator, dual-track (dense-tag 57 / structured bilingual 22)
ProvenanceField-collected photographs and controlled generation, manually screened and deduplicated
Public samples12 of 79 pairs published on this page

Citing & licensing

The dataset page on the lab site carries the current license text and version history. If you build on it, tell us — we keep a running list of derived work.