We are sunsetting our office from Canada and shifting to Hong Kong. Stay connected for latest updates; physical meetings will resume soon.
Pricing Plans

Transparent Pricing Models

Pay for structured data volume. Choose a pre-built plan or adjust our dynamic token slider to fit your exact training budget.

Standard Ingestion

Standard

Standard web dataset compilation for pretraining 7B-13B parameter models.

$30,000/ contract
  • 500M clean tokens
  • 16 concurrent crawling threads
  • Rotating residential proxy exits
  • Hugging Face & pgvector sync (manual)
Dynamic Volume Plan
Scale Scraper

Scale Ingest

Adjust slider to compile larger dataset partitions.

$50,000/ contract
Ingestion Tokens:10B tokens / contract
Worker Concurrency:16 Concurrency Threads
Residential Proxies:1,500 Residential Nodes
Base Ingestion Fee:$35,000
Proxy Routing Allocation:$15,000
  • Auto PII scrub & deduplication
  • Hugging Face & Pinecone export sync
  • High-speed residential rotating proxies
Bespoke Pretraining

Foundation

Bespoke multi-trillion token datasets and dedicated regional crawling clusters.

$2,000,000+/ contract
  • 1T+ high-quality pretraining tokens
  • Dedicated exit proxy subnets (unlimited)
  • Bespoke parsers & 24/7 Slack engineers
  • Real-time AWS S3 / Redshift streams

Technical Specification Matrix

Crawler ParameterStandardScale IngestFoundation
Ingestion Dataset Volume500M tokens10B - 100B tokens1T+ custom corpus
Concurrent Crawler Threads16 threadsUp to 256 threadsDedicated Cluster
IP Residential Proxy PoolShared DatacenterRotating Residential (15k+)Dedicated Whitelist Nodes
MinHash LSH DeduplicationStandard cleanFully automated v2Custom Jaccard tuning
PII Anonymization scrubberBasic regexAutomated regex/NLPCustom Entity classifiers
Destination Sync TargetsManual DownloadHugging Face, Milvus, pgvectorAWS S3, Redshift stream
CAPTCHA solver integrationNot availableCNN solve model inlineBespoke solve routing
Support Response LatencyEmail (48h)Slack & Email (<12h)Dedicated solutions engineer

Pricing FAQ

How are tokens calculated?

One token corresponds to approximately 4 characters of parsed markdown text or raw structured JSON outputs. Repetitive boilerplate code stripped by the LSH sweep does not count towards token usage.

Is there a daily scraping limit?

No. Your concurrent thread capacity determines how fast data is compiled. Free plan users are capped at 2 concurrency threads.

Can I bring my own proxies?

Yes. In the Scale plan and above, you can specify custom residential proxies. By default, Reworkd manages exit nodes automatically.

What formats do you support?

JSON, JSONL, Parquet, and CSV file formats. You can sync directly to Hugging Face, S3 buckets, or Pinecone databases.