Skip to content

data infrastructure

Training datathat carries itssource.

PEB provides four datasets - Textbook, Coding, 3D CAD, and Infographics - each one delivered with its record-level schema, provenance, and license posture attached, not asserted after the fact.

~2.5MRecords across four datasets
Record-level provenanceBuilt-in trust
License posture attachedNot asserted after the fact
Six-stage quality pipelineApplied to every record

PEB provides AI training data that is attributable to its source and checkable at the record level.

For the evaluating engineer

Compare all four datasets.

Modality, formats, attribution, license posture, delivery, and typical training stage - the single most useful thing here for technical review.

Comparison of all four datasets across modality, formats, attribution, licensing, delivery, and training stage
 Textbook DatasetCoding Dataset3D CAD DatasetInfographics Dataset
ModalityTextCode3D geometryImage + structured data
Primary formatsJSONL, ParquetJSONL, Parquet, ArrowSTEP, glTF, OBJPNG, SVG, JSON
Attribution basisISBN + publisher catalogRepository + license recordPrimary domain + source typeCategory + rendering provenance
License postureCustom license, per engagementPer-repository (Legacy) / custom (DSA)Custom license, per engagementCustom license, per engagement
DeliveryS3 / private HF repoS3 / private HF repoS3 / GCS bucketS3 / GCS bucket
Typical training stagePre-training, reasoning SFTPre-training, code SFTPre-training, spatial & multimodal SFTPre-training, document & vision-language SFT

Scroll to see all columns

What happens before delivery

Six stages, applied to every record.

Acquisition through human audit sampling - the same sequence runs across all four modalities. See the full methodology for how each dataset applies it.

  1. Acquisition

    Sources material from publishers, repositories, CAD archives, and chart generation.

  2. Rights verification

    Confirms a licensing or ownership basis before intake.

    Removes material without a confirmed rights basis.

  3. Normalization

    Converts source formats into each modality's delivery schema.

    Removes files that fail to parse or normalize cleanly.

  4. Deduplication

    Compares hashes and near-duplicates per modality.

    Removes redundant and near-identical records.

  5. Contamination screening

    Checks against named public benchmark test splits, per modality.

    Removes records that overlap a benchmark test or validation split.

  6. Human audit sampling

    Manually reviews a sampled subset before a release is tagged.

    Removes records that fail manual review.

The proof, not the pitch

A real record from each dataset.

Rounded and redacted only where the record itself required it. Copy any of them and check them against the schema on that dataset's page.

Textbook Dataset · 1 shown
[  {    "isbn": "9788169950817",    "title": ".NET Application Development Lab",    "page_count": 161,    "word_count": 18345,    "category": "Computer Science & IT",    "stem_flag": true,    "language": "English"  }]
Coding Dataset · 1 shown
[  {    "problem_title": "Shortest Path in a Binary Weight Graph",    "topic": "DSA",    "languages": [      "cpp",      "java",      "python",      "csharp",      "javascript"    ],    "cpp_lines": 67,    "cpp_tokens": 535  }]
3D CAD Dataset · 1 shown
[  {    "primary_domain": "Architectural Engineering",    "primary_category": "Engineering",    "dataset_count": 378772,    "distribution_pct": 28.79  }]
Infographics Dataset · 1 shown
[  {    "main_category": "Engineering, Technology & Infrastructure",    "subcategory": "IT, Software & Cloud Computing",    "image_count": 46500  }]

Delivery and formats

One access model across all four datasets. Per-dataset format detail lives on each dataset's own page.

Formats
JSONL / Parquet / Arrow (text, code) · STEP / glTF / OBJ (geometry) · PNG / SVG / JSON (infographics)
Mechanisms
S3 pre-signed URL, GCS bucket, private Hugging Face repository
Versioning
Immutable dated snapshots, referenced by release tag
Update cadence
Quarterly refresh across all four datasets
Access model
Evaluation copy on request; full access under a signed data agreement

Where the data comes from

PEB Educational Services began as an educational publisher and curriculum developer, producing and licensing textbook content across STEM and non-STEM subjects for schools and institutions.

That publishing relationship is the basis for the Textbook Dataset’s ISBN-level attribution - every record traces back to a title PEB already held rights to catalog, not material assembled after the fact.

Read about the education arm

Read this before you evaluate us

What we don’t claim.

  • We do not claim comprehensive language coverage - a handful of languages account for the majority of Textbook Dataset word volume, and the rest remain a minority share.
  • We do not measure or publish benchmark accuracy uplift from training on this data - contamination screening confirms overlap removal, not downstream performance gain.
  • We do not warrant fitness for safety-critical or medical use of the 3D CAD Dataset - geometry accuracy has not been certified to an engineering standard.
  • We do not offer real-time or continuously streaming delivery - every dataset ships as a dated, versioned snapshot, refreshed quarterly.

Tell us what you’re training and we’ll send an evaluation copy.

Sample records first, full access under a signed data agreement where the use fits.