Privacy Schema Agreement
Last Updated: July 4, 2026. This Privacy Schema outlines our policies regarding data harvesting, PII redaction compliance, and target dataset distribution.
1. Data Collection Philosophy
Reworkd AI builds automated high-throughput web scraping workers designed to compile public web text into structured training datasets for generative AI models. We process data strictly for machine learning pretraining, semantic lookup, and alignment tuning.
We operate exclusively on public domain resources, open sitemaps, and websites that do not prohibit automated crawling headers. We do not crawl authenticated, password-protected portals or private user areas.
2. PII Scrubbing and Filtering
All harvested raw text blocks undergo immediate, automated pipeline processing before they are stored or compiled into Parquet files. Our sanitization pipelines apply:
- Named Entity Recognition (NER) models to mask personal identifiers (e.g. human names).
- Regex patterns targeting emails, telephone paths, physical addresses, and security keys.
- Hash signature filters to strip user session metadata variables.
This ensures that the final dataset contains only general semantic concepts, preventing pretraining models from memorizing private customer coordinates.
3. Deletion and Right-to-Forgot
In accordance with GDPR, CCPA, and global data commons principles, Reworkd AI honors opt-out and deletion requests. Website publishers or individuals who discover sensitive information inside our published indices can request target redaction.
Once a deletion request is validated, Reworkd AI removes the corresponding indices from our active vector pools, updates the lineage index logs, and reflects the scrubbing update downstream in client-synchronized datasets.