data infrastructure
Training datathat carries itssource.
PEB provides four datasets - Textbook, Coding, 3D CAD, and Infographics - each one delivered with its record-level schema, provenance, and license posture attached, not asserted after the fact.
The catalog
Four datasets, one delivery model.
Four datasets, one register
- Textbook DatasetEducational text carrying its ISBN, edition category, and page span, drawn from academic and institutional publishers across fifteen languages.JSONL / Parquet10K–100K recordsCustom license
- Coding DatasetTwo code sub-collections: independently authored DSA problem-solution pairs across seven languages, and production Legacy Codebases with full git history across fourteen industries.JSONL / Parquet / Arrow10K–100K recordsPer-repo / custom license
- 3D CAD DatasetParametric CAD and mesh geometry indexed by primary domain, spanning engineering, biology, culture, and twenty other domains, each record carrying source type and license basis.STEP / glTF / OBJ1M+ recordsCustom license
- Infographics DatasetChart and diagram images paired with the structured data used to render them, spanning seven subject categories from natural sciences to business and finance.PNG / SVG / JSON1M+ recordsCustom license
PEB provides AI training data that is attributable to its source and checkable at the record level.
For the evaluating engineer
Compare all four datasets.
Modality, formats, attribution, license posture, delivery, and typical training stage - the single most useful thing here for technical review.
| Textbook Dataset | Coding Dataset | 3D CAD Dataset | Infographics Dataset | |
|---|---|---|---|---|
| Modality | Text | Code | 3D geometry | Image + structured data |
| Primary formats | JSONL, Parquet | JSONL, Parquet, Arrow | STEP, glTF, OBJ | PNG, SVG, JSON |
| Attribution basis | ISBN + publisher catalog | Repository + license record | Primary domain + source type | Category + rendering provenance |
| License posture | Custom license, per engagement | Per-repository (Legacy) / custom (DSA) | Custom license, per engagement | Custom license, per engagement |
| Delivery | S3 / private HF repo | S3 / private HF repo | S3 / GCS bucket | S3 / GCS bucket |
| Typical training stage | Pre-training, reasoning SFT | Pre-training, code SFT | Pre-training, spatial & multimodal SFT | Pre-training, document & vision-language SFT |
Scroll to see all columns
What happens before delivery
Six stages, applied to every record.
Acquisition through human audit sampling - the same sequence runs across all four modalities. See the full methodology for how each dataset applies it.
Acquisition
Sources material from publishers, repositories, CAD archives, and chart generation.
Rights verification
Confirms a licensing or ownership basis before intake.
Removes material without a confirmed rights basis.
Normalization
Converts source formats into each modality's delivery schema.
Removes files that fail to parse or normalize cleanly.
Deduplication
Compares hashes and near-duplicates per modality.
Removes redundant and near-identical records.
Contamination screening
Checks against named public benchmark test splits, per modality.
Removes records that overlap a benchmark test or validation split.
Human audit sampling
Manually reviews a sampled subset before a release is tagged.
Removes records that fail manual review.
The proof, not the pitch
A real record from each dataset.
Rounded and redacted only where the record itself required it. Copy any of them and check them against the schema on that dataset's page.
[ { "isbn": "9788169950817", "title": ".NET Application Development Lab", "page_count": 161, "word_count": 18345, "category": "Computer Science & IT", "stem_flag": true, "language": "English" }][ { "problem_title": "Shortest Path in a Binary Weight Graph", "topic": "DSA", "languages": [ "cpp", "java", "python", "csharp", "javascript" ], "cpp_lines": 67, "cpp_tokens": 535 }][ { "primary_domain": "Architectural Engineering", "primary_category": "Engineering", "dataset_count": 378772, "distribution_pct": 28.79 }][ { "main_category": "Engineering, Technology & Infrastructure", "subcategory": "IT, Software & Cloud Computing", "image_count": 46500 }]Delivery and formats
One access model across all four datasets. Per-dataset format detail lives on each dataset's own page.
- Formats
- JSONL / Parquet / Arrow (text, code) · STEP / glTF / OBJ (geometry) · PNG / SVG / JSON (infographics)
- Mechanisms
- S3 pre-signed URL, GCS bucket, private Hugging Face repository
- Versioning
- Immutable dated snapshots, referenced by release tag
- Update cadence
- Quarterly refresh across all four datasets
- Access model
- Evaluation copy on request; full access under a signed data agreement
Where the data comes from
PEB Educational Services began as an educational publisher and curriculum developer, producing and licensing textbook content across STEM and non-STEM subjects for schools and institutions.
That publishing relationship is the basis for the Textbook Dataset’s ISBN-level attribution - every record traces back to a title PEB already held rights to catalog, not material assembled after the fact.
Read this before you evaluate us
What we don’t claim.
- We do not claim comprehensive language coverage - a handful of languages account for the majority of Textbook Dataset word volume, and the rest remain a minority share.
- We do not measure or publish benchmark accuracy uplift from training on this data - contamination screening confirms overlap removal, not downstream performance gain.
- We do not warrant fitness for safety-critical or medical use of the 3D CAD Dataset - geometry accuracy has not been certified to an engineering standard.
- We do not offer real-time or continuously streaming delivery - every dataset ships as a dated, versioned snapshot, refreshed quarterly.
Tell us what you’re training and we’ll send an evaluation copy.
Sample records first, full access under a signed data agreement where the use fits.