A data factory for teams training AI models

Source data, clean it, generate missing records, label what matters, and export reviewed Data Packs from one workspace. Not just annotation. Not just synthetic. Dataset operations.

Find
Acquire
Profile
Clean
Generate / Label
Export
daqa / search / invoice-receipts

Ask Daqa

“I need invoice and receipt data for field extraction”

Sources checked · 3 sources · 10 result links
  • Hugging Face Datasets

    api / needs_preparation

    Found 5
  • Kaggle Datasets

    api / needs_preparation

    Found 2
  • Zenodo

    api / research_ready

    Found 3
Needs inspect
Direct leads
Top candidates
  • Francisco-Cruz/InvoicesReceiptsPT

    Hugging Face

    Use this

    Apache-2.0 / needs_preparation / allowed

  • Dataset of invoices and receipts including annotation of relevant fields

    Zenodo

    Use this

    CC-BY / needs_preparation / allowed

  • Dataset of personal invoices and receipts including annotation of relevant fields

    Zenodo

    Use this

    CC-BY / needs_preparation / allowed

Field targets
6 fields
vendordatetotaltaxline_itemsbounding_boxes
Export target

JSONL + document images + provenance

Every step of building training data, in one place

Scan the workspace, not a wall of text. Each capability is designed to be understood at a glance and connected to the same dataset, schema, and provenance.

Data sourcing

Find useful public datasets across providers, inspect external results, and import candidates instead of manually searching platform by platform.

KaggleHugging FaceGitHubZenodoOpenMLRoboflowFigshareDryadNASA CMRData.govand more

Acquire & upload

Import a candidate or upload your own files. Daqa tracks where each dataset came from and what still needs action.

  • web sourceselected
  • local filesuploaded
  • external leadneeds action

Profile & define schema

Detect columns, missing values, duplicates, and likely labels, then approve the schema, roles, and sensitive fields.

idint · unique
imageasset · 0 null
labelenum · 3%

Cleaning & transformation you approve

Daqa proposes steps like dropping useless fields, normalizing labels, handling missing values, and shaping records, but nothing is final until you approve it.

Drop empty columns
Normalize labels
Impute missing

Synthetic generation

Plan and review synthetic data for missing or underrepresented cases. Outputs keep model, prompt, seed, review state, and provenance.

modelpromptseedreviewprovenance

Native labeling

Define label schemas and review annotations for text, tabular, and image records connected to the pipeline, provenance, and export.

positiveneutralnegativespam

Auto-labeling with review

Models suggest labels with confidence and provenance. They're suggestions, not trusted labels. You accept or reject.

invoice · 0.93

CVAT for advanced vision

Optional CVAT backend for boxes, polygons, masks, keypoints, video tracking, and 3D. Daqa owns the dataset, schema, and export.

boxespolygonsmaskskeypointsvideo tracking3D

Bounded web collection

Search the web, inspect URLs, crawl a bounded site map, extract candidate records, and review license and risk signals before import.

→ crawl: docs.example.com/*

✓ candidate records extracted

⚠ license signals to review

Discover where relevant data may live

Provider cards surface external catalog leads across public data platforms. External results aren't treated as Daqa-owned or verified until you import, profile, and review them.

Every source carries signals

License evidence
Provenance trail
Validation notes
Risk signals

Daqa shows source context and risk. It isn't a legal-clearance product.

K

Kaggle

Dataset hub

External search

H

Hugging Face

Datasets & models

External search

G

GitHub

Repos & releases

Source leads

Z

Zenodo

Research data

External search

O

OpenML

ML benchmarks

Catalog source

+

And more

Public data providers

Expanding catalog

Every dataset runs through the same AI-readiness pipeline

No matter where data comes from, it becomes usable through one repeatable workflow, not one-off scripts. Decisions stay reviewable at every stage.

01

Import

Public, uploaded, web-collected, or generated.

02

Profile

Columns, assets, missing values, risk signals.

03

Define schema

Approve labels, roles, sensitive fields.

04

Clean

Drop, normalize, and fix. You approve each step.

05

Transform

Shape records into training-ready form.

06

Validate

Checks, notes, and quality gates.

07

Dataset card

Provenance, license, and generation metadata.

08

Export

Model-ready Data Pack in the right format.

Export model-ready Data Packs

The final output isn't just a cleaned-looking table. A Data Pack bundles everything you need to actually train, with the context to trust what you're training on.

Files & records
Schema
Labels
Generation metadata
License & risk evidence
Validation & provenance
example-datapack.manifest

dataset export manifest

records, files, schema, decisions

reviewable

Format targets

JSONLCSVImage manifestCOCOYOLOPreference pairs

provenance: source → profile → approved steps → export

license: source evidence preserved for review

validation: checks, warnings, and notes included

Workspace plans, plus usage credits

Plans include hard usage limits plus AI credits. Static crawling and deterministic pipeline work use quotas, while expensive AI extraction, generation, labeling, and media analysis use credits.

Free

100 AI credits/mo

Try Daqa

$0

Preview dataset search and run tiny pipeline jobs before upgrading.

1 user1 private project1 GB storage5k records/mo100 crawl pages/mo
  • Dataset search previews
  • Tiny imports and profiling
  • 2 sample exports
Join waitlist

Builder

2,000 AI credits/mo

Solo dataset work

$30/mo

$24/mo billed yearly

Solo AI builders preparing small real datasets with AI search and bounded crawling.

1 user3 private projects10 GB storage50k records/mo2,500 crawl pages/mo
  • 15 pipeline jobs
  • 10 Data Pack exports
  • 500 MB max upload
Join waitlist

Pro

3,500 AI credits/mo

Serious solo workflows

$89/mo

$72/mo billed yearly

Freelancers, consultants, and serious solo users who outgrow Builder.

1 user10 private projects50 GB storage250k records/mo10k crawl pages/mo
  • 50 pipeline jobs
  • Limited auto-labeling
  • Light Developer API
Join waitlist

Enterprise

Custom AI credits

Custom & on-prem

Custom

Private cloud, self-hosting, regulated data, and custom workflows.

SSO/SAMLPrivate storagePrivate LLM routingCustom limits
  • Dedicated infrastructure
  • Security review support
  • Custom connectors
Contact us

Usage credits

Each plan includes a monthly AI credit allowance. Credits are spent only when Daqa asks AI to do meaningful work, such as search planning, page inspection, record extraction, generation, auto-labeling, media analysis, or dataset card writing.

Uploading files, browsing sources, static crawling within your page quota, deterministic profiling, approved cleaning steps, and normal exports use plan limits instead of AI credits. Larger jobs show an estimated credit cost before they run, and overages stay opt-in.

  • Included credits reset every month.
  • Small chat clarifications stay lightweight so the product feels natural.
  • Heavy AI extraction, generation, and auto-labeling are estimated before running.
Normal AI search3 credits
Dataset discovery run10 credits
Deep discovery25 credits
Pipeline advisor15 credits
Dataset card40 credits
AI inspect 10 URLs10 credits
AI rank 100 pages50 credits
AI extract 100 pages200 credits
Small generation preview100 credits
Generate 1k text records400 to 600 credits

Get early access to the data factory

Tell us how you build training data today. We're onboarding teams gradually and shaping Daqa around the workflows that matter most.

One workspace

Public, web-collected, generated, labeled, and uploaded data in one place.

Reviewable by design

Decisions stay traceable. Provenance, license, and validation travel with the data.

Model-ready output

Export Data Packs your team can actually train on, not just a cleaned table.

Workflows that matter

Pick any that apply.

Dataset types

No spam. We'll only email about early access.