We are sunsetting our office from Canada and shifting to Hong Kong. Stay connected for latest updates; physical meetings will resume soon.
AI Dataset Engine v2.5Active

Ingest raw web data at scale,formatted for AI models.

Automatically crawl, parse, deduplicate, and vectorize web targets into pristine training datasets. Direct pipelines to Hugging Face, Pinecone, and Parquet.

96%
Proxy Health
Rotating Nodes
89%
Boilerplate Sweep
LSH Deduplication
99%
API SLA Buffer
99.9% Uptime
Operations Panel

Ingestion Orchestrator

Control target scraping speeds, re-route proxy region configurations, and preview raw-to-JSON extractions.

Ingestion Dashboard
Scraper Threads: 16
Region:

Systems load

Scraper CPU43%
Buffer Pool47%
Scrape statistics
Throughput:6,880 tokens/s
Proxy Count:32 Nodes
Active Pipelines4 targets online
arXiv AI Corpus
arxiv.org
8.4M tokensActive
TechCrunch Feeds
techcrunch.com
2.1M tokensActive
GitHub JS Trends
github.com
14.2M tokensActive
Finance Index Feed
bloomberg.com
0 tokensIdle
DEDUP ALGORITHM: MINHASH LSHSYNC DEST: HUGGINGFACE

Live Parser Parser

Scraped DOM (github.com/trending)
<body><h2>Trending repositories</h2><div class='repo'>react - Declarative UI library (240k stars)</div></body>
Structured JSON OutputPROCESSED
{
  "source": "github.com",
  "title": "Trending repositories",
  "items": [
    {"name": "react", "stars": 240000}
  ],
  "tokens": 481
}
Initialized scraper worker pool successfully
Connected proxy nodes to USA exit region
Linked ingestion pipeline targets [DS-109]
Ingestion Flow Wave:
Product Features

Designed for AI Engineering

A comprehensive pipeline stack for data vectorization, schema validation, and strict compliancy scraping.

Builder Module

Interactive Schema Modeler

Specify and parse custom database parameters on the fly. Click fields below to customize the JSON schema output dynamically.

{
  "dataset_schema": {
    "source": "url_string",
    "title": "string_parsed",
    "content": "markdown_clean",
    "vector_embedding": "float32[1536]"
  }
}
Vector Sync

Semantic Vector Chunking

Slices parsed raw structures into text nodes, runs embed algorithms, and deposits vectors dynamically.

CHUNK SIZE: 512BUPDATING
0
1
2
3
4
5
6
7
8
9
Parser Cleaning

LSH Deduplication

Removes repetitive banners, duplicate site footers, and template noise, preserving raw content data.

-header list items
-newsletter form container
+clean parsed paragraphs
Compliance

PII Scrubber & Audit Log

Enforces robots.txt guidelines and automatically scrubs emails, phone numbers, and location values before compilation.

File ScannedPII Statusrobots.txt
arxiv_2402_14.pdfCLEANPASSED
github_readme.mdCLEANPASSED
techcrunch_feed.htmlSCRUBBEDPASSED
Dataset Registries

Curated Dataset Gallery

Browse structured datasets ready for download, fine-tuning, and vector mapping.

AVAILABLE SHARDS (3)

Reddit Dialogue & Opinion Graph

28 Billion tokens

Anonymized conversational trees from top informational subreddits, clean of bot spam, structured for training LLM conversational agents.

FORMAT: JSONL (Gzipped)LICENSE: Reddit Public API Compliance

Facebook Public Page Marketing Corpora

14 Billion tokens

Extracted post data from verified corporate and news public pages. Optimized for marketing copywriting and brand sentiment models.

FORMAT: ParquetLICENSE: Public Media Lineage

AI Academic Papers Corpus

42 Billion tokens

Pruned academic PDFs parsed into structured markdown blocks containing titles, abstracts, and reference annotations.

FORMAT: Parquet (Sharded)LICENSE: OpenAccess (CC-BY-4.0)
DATASET ROW SHARD PREVIEW
PREVIEW REGISTRY

Reddit Dialogue & Opinion Graph

Anonymized conversational trees from top informational subreddits, clean of bot spam, structured for training LLM conversational agents.

{
  "subreddit": "r/machinelearning",
  "thread_id": "t3_1dx492a",
  "discussion_text": "Evaluating local model performance on complex logic...",
  "replies": [
    {"author_karma": 12800, "body": "Llama 3 8B scales well if fine-tuned on clean datasets..."}
  ],
  "upvotes": 412
}
User Portfolio

Powering Boutique AI Labs

Meet the specialized, niche AI research teams and startups training their models on reworkd structured datasets.

4.2B tokens

DeepSoil AI

Crop-Yield Modeling & Agritech

Trains local soil humidity and crop-yield prediction models on localized agricultural web records and historical weather datasets.

STREAM ACTIVE
8.5B tokens

Lexicon Legal Tech

Contract Summary NLP

Refines legal text summarization and contract compliance algorithms using structured public court transcripts and legal filing archives.

STREAM ACTIVE
6.8B tokens

Aetheris Research

Ecological Weather Forecasts

Develops niche ecological simulation systems using environmental web archives, monitoring global sensor registers.

STREAM ACTIVE
5.4B tokens

VoxelCraft AI

Indie Game Texturing

Trains game texture synthesis and text-to-asset models using parsed public catalog metadata and image alignments.

STREAM ACTIVE
3.1B tokens

OmniScribe NLP

Medical Speech-To-Text

Trains medical conversation transcription systems on specialized regional clinic publications and public files.

STREAM ACTIVE
2.7B tokens

Verba Translation

Minority Language Dialects

Fine-tunes minority language translation engines on regional conversational and public dialogue logs.

STREAM ACTIVE
Ingestion Flow

Raw-To-Train Pipeline

Six distinct pipeline stages that sanitize, vectorise, and sharded web nodes for model pretraining.

Stage 01

URL Discovery

Smart crawlers map targets, parse domain sitemaps, resolve CAPTCHAs, and index outbound links into the queues dynamically.

Stage 02

DOM Extraction

Strips layouts of CSS stylesheets, tracking pixels, scripts, and sidebar links, extracting clean semantic Markdown.

Stage 03

PII Clean-Up

An automated regex and NLP sanitization loop scrubs emails, phone numbers, addresses, and handles robots.txt compliance.

Stage 04

LSH Deduplication

Executes MinHash LSH checks at scale, filtering repetitive paragraphs and duplicate files to prevent LLM training overfit.

Stage 05

Vector Chunking

Tokenizes string text by Llama or OpenAI token indexes, segments into vectors, and generates float32 database embeddings.

Stage 06

Parquet Export

Compresses processed tokens into sharded Parquet files and pushes datasets directly to Hugging Face or AWS S3 pools.

480B+
Tokens Processed
99.94%
Data Integrity Ratio
1.5K+
Active Pipelines
60+
Connected Repositories
Model Ingestion

Aligned for Generative Training.

Reworkd datasets compile and map web text structures directly into the exact instruction wrappers required by state-of-the-art models. Skip formatting scripts and train instantly.

Llama 3 & Mistral Instruction AlignmentsPre-Formatted

Datasets are mapped to raw instruction-tuning wrappers (ChatML tags) to facilitate pre-packaged SFT (Supervised Fine-Tuning) paths directly.

<|begin_of_text|><|start_header_id|>system<|end_header_id|>
You are a helpful assistant.
<|start_header_id|>user<|end_header_id|>
{{instruction_crawled_data}}
<|start_header_id|>assistant<|end_header_id|>

Start Ingesting Datasets

Spin up scraper worker clusters instantly. Connect raw schema variables and stream dataset blocks straight to your training pipeline.