A data factory for teams training AI models
Source data, clean it, generate missing records, label what matters, and export reviewed Data Packs from one workspace. Not just annotation. Not just synthetic. Dataset operations.
Ask Daqa
“I need invoice and receipt data for field extraction”
- Found 5
Hugging Face Datasets
api / needs_preparation
- Found 2
Kaggle Datasets
api / needs_preparation
- Found 3
Zenodo
api / research_ready
- Use this
Francisco-Cruz/InvoicesReceiptsPT
Hugging Face
Apache-2.0 / needs_preparation / allowed
- Use this
Dataset of invoices and receipts including annotation of relevant fields
Zenodo
CC-BY / needs_preparation / allowed
- Use this
Dataset of personal invoices and receipts including annotation of relevant fields
Zenodo
CC-BY / needs_preparation / allowed
JSONL + document images + provenance
Every step of building training data, in one place
Scan the workspace, not a wall of text. Each capability is designed to be understood at a glance and connected to the same dataset, schema, and provenance.
Data sourcing
Find useful public datasets across providers, inspect external results, and import candidates instead of manually searching platform by platform.
Acquire & upload
Import a candidate or upload your own files. Daqa tracks where each dataset came from and what still needs action.
- web sourceselected
- local filesuploaded
- external leadneeds action
Profile & define schema
Detect columns, missing values, duplicates, and likely labels, then approve the schema, roles, and sensitive fields.
Cleaning & transformation you approve
Daqa proposes steps like dropping useless fields, normalizing labels, handling missing values, and shaping records, but nothing is final until you approve it.
Synthetic generation
Plan and review synthetic data for missing or underrepresented cases. Outputs keep model, prompt, seed, review state, and provenance.
Native labeling
Define label schemas and review annotations for text, tabular, and image records connected to the pipeline, provenance, and export.
Auto-labeling with review
Models suggest labels with confidence and provenance. They're suggestions, not trusted labels. You accept or reject.
CVAT for advanced vision
Optional CVAT backend for boxes, polygons, masks, keypoints, video tracking, and 3D. Daqa owns the dataset, schema, and export.
Bounded web collection
Search the web, inspect URLs, crawl a bounded site map, extract candidate records, and review license and risk signals before import.
→ crawl: docs.example.com/*
✓ candidate records extracted
⚠ license signals to review
Discover where relevant data may live
Provider cards surface external catalog leads across public data platforms. External results aren't treated as Daqa-owned or verified until you import, profile, and review them.
Every source carries signals
Daqa shows source context and risk. It isn't a legal-clearance product.
Kaggle
Dataset hub
External search
Hugging Face
Datasets & models
External search
GitHub
Repos & releases
Source leads
Zenodo
Research data
External search
OpenML
ML benchmarks
Catalog source
And more
Public data providers
Expanding catalog
Every dataset runs through the same AI-readiness pipeline
No matter where data comes from, it becomes usable through one repeatable workflow, not one-off scripts. Decisions stay reviewable at every stage.
Import
Public, uploaded, web-collected, or generated.
Profile
Columns, assets, missing values, risk signals.
Define schema
Approve labels, roles, sensitive fields.
Clean
Drop, normalize, and fix. You approve each step.
Transform
Shape records into training-ready form.
Validate
Checks, notes, and quality gates.
Dataset card
Provenance, license, and generation metadata.
Export
Model-ready Data Pack in the right format.
Export model-ready Data Packs
The final output isn't just a cleaned-looking table. A Data Pack bundles everything you need to actually train, with the context to trust what you're training on.
dataset export manifest
records, files, schema, decisions
Format targets
provenance: source → profile → approved steps → export
license: source evidence preserved for review
validation: checks, warnings, and notes included
Workspace plans, plus usage credits
Plans include hard usage limits plus AI credits. Static crawling and deterministic pipeline work use quotas, while expensive AI extraction, generation, labeling, and media analysis use credits.
Usage credits
Each plan includes a monthly AI credit allowance. Credits are spent only when Daqa asks AI to do meaningful work, such as search planning, page inspection, record extraction, generation, auto-labeling, media analysis, or dataset card writing.
Uploading files, browsing sources, static crawling within your page quota, deterministic profiling, approved cleaning steps, and normal exports use plan limits instead of AI credits. Larger jobs show an estimated credit cost before they run, and overages stay opt-in.
- Included credits reset every month.
- Small chat clarifications stay lightweight so the product feels natural.
- Heavy AI extraction, generation, and auto-labeling are estimated before running.
Get early access to the data factory
Tell us how you build training data today. We're onboarding teams gradually and shaping Daqa around the workflows that matter most.
One workspace
Public, web-collected, generated, labeled, and uploaded data in one place.
Reviewable by design
Decisions stay traceable. Provenance, license, and validation travel with the data.
Model-ready output
Export Data Packs your team can actually train on, not just a cleaned table.