Transparent Pricing Models
Pay for structured data volume. Choose a pre-built plan or adjust our dynamic token slider to fit your exact training budget.
Standard
Standard web dataset compilation for pretraining 7B-13B parameter models.
- 500M clean tokens
- 16 concurrent crawling threads
- Rotating residential proxy exits
- Hugging Face & pgvector sync (manual)
Scale Ingest
Adjust slider to compile larger dataset partitions.
- Auto PII scrub & deduplication
- Hugging Face & Pinecone export sync
- High-speed residential rotating proxies
Technical Specification Matrix
| Crawler Parameter | Standard | Scale Ingest | Foundation |
|---|---|---|---|
| Ingestion Dataset Volume | 500M tokens | 10B - 100B tokens | 1T+ custom corpus |
| Concurrent Crawler Threads | 16 threads | Up to 256 threads | Dedicated Cluster |
| IP Residential Proxy Pool | Shared Datacenter | Rotating Residential (15k+) | Dedicated Whitelist Nodes |
| MinHash LSH Deduplication | Standard clean | Fully automated v2 | Custom Jaccard tuning |
| PII Anonymization scrubber | Basic regex | Automated regex/NLP | Custom Entity classifiers |
| Destination Sync Targets | Manual Download | Hugging Face, Milvus, pgvector | AWS S3, Redshift stream |
| CAPTCHA solver integration | Not available | CNN solve model inline | Bespoke solve routing |
| Support Response Latency | Email (48h) | Slack & Email (<12h) | Dedicated solutions engineer |
Pricing FAQ
How are tokens calculated?
One token corresponds to approximately 4 characters of parsed markdown text or raw structured JSON outputs. Repetitive boilerplate code stripped by the LSH sweep does not count towards token usage.
Is there a daily scraping limit?
No. Your concurrent thread capacity determines how fast data is compiled. Free plan users are capped at 2 concurrency threads.
Can I bring my own proxies?
Yes. In the Scale plan and above, you can specify custom residential proxies. By default, Reworkd manages exit nodes automatically.
What formats do you support?
JSON, JSONL, Parquet, and CSV file formats. You can sync directly to Hugging Face, S3 buckets, or Pinecone databases.