Expert Curated Data Samples
Explore high-quality datasets across key domains, purpose-built for real-world model development. Choose from ready-to-use packs or share a request for custom data solutions.
Latest Additions

New
Physical Intelligence
Structured, real-world datasets for robotic perception, action understanding, and environment-awaredecision-making.
.png)
New
RL as Rubrics
Labeled taxonomy of eighteen rubric-writing skills across five categories and three proficiency levels for competency mapping and annotator calibration.
.png)
New
Multi-image reference to Video generation
Multi-reference images and edit prompts paired with human-scored 10-second videos assessed across six dimensions for video-generation evaluation.
%20(1).png)
Dataset Library
Explore curated data collections designed to accelerate model development across training, alignment, and evaluation workflows.
1 datapoints
Image
video
Generalist
Vision Annotation
Images, video, and LiDAR point clouds annotated across nine types including bounding boxes, segmentation, and 3D cuboids for perception models.
View dataset ↗
70 datapoints
Generalist
SFT
Video
Motion capture video samples
Seventy synchronized motion-capture samples pairing skeleton overlays, 3D renders, and egocentric video across household tasks for pose estimation.
View dataset ↗
30 datapoints
Video
Generalist
SFT
Egocentric Hand Keypoint Annotation
First-person household video with 21-point 3D hand keypoints across cooking, cleaning, and other daily tasks for hand-pose estimation.
View dataset ↗
42 datapoints
Video
Generalist
SFT
Egocentric video data
Twenty narrated first-person recordings of household and industrial tasks with think-aloud commentary for embodied action understanding.
View dataset ↗
5 datapoints
Video
Generalist
SFT
MultiCam
Fifteen synchronized head- and wrist-camera recordings of manipulation tasks like clipping and threading for bimanual policy training.
View dataset ↗
39 datapoints
Video
Generalist
SFT
Egocentric household annotated
Fifteen first-person household recordings with fine-grained hand and grasp annotations across domestic tasks for action recognition training.
View dataset ↗
20 datapoints
Video
Generalist
SFT
Egocentric factory
Sixteen first-person clips of electronics assembly and PCB testing on an active production line for industrial task recognition.
View dataset ↗
3 datapoints
Video
Generalist
SFT
Egocentric Factory De-IDed
Face-blurred first-person production-line video showing PCB handling and record-keeping for de-identified industrial action recognition.
View dataset ↗
20 datapoints
RLHF
Text
Experts - Coding
Preference Pair Trajectory
Passing and failing coding-agent trajectories from matched Terminal Bench 2 tasks, paired for preference optimization and debugging-agent training.
View dataset ↗
15 datapoints
Eval
Operations
TEXT
Terminal bench tasks
Ten executable terminal tasks spanning workflows including data-pipeline debugging, systems programming, and security analysis, with deterministic verifiers for evaluation.
View dataset ↗
10 datapoints
Eval
Generalist
TEXT
Cultural Reasoning
Native-expert prompts requiring local cultural knowledge, scored across 11 dimensions to evaluate multilingual model reasoning.
View dataset ↗
14 datapoints
RLHF
Language Experts
TEXT
RLHF
Indic-language prompts with four ranked model responses and rater explanations across six dimensions for preference-model training.
View dataset ↗
10 datapoints
Eval
Language Experts
TEXT
SFT Image + Text
Document, chart, receipt, table, and diagram images with ranked model responses and golden answers for vision-language model fine-tuning.
View dataset ↗
9 datapoints
Eval
Language Experts
TEXT
Linguistic Reasoning
Native-expert prompts testing grammar, morphology, syntax, semantics, and usage, scored across 11 dimensions for multilingual reasoning evaluation.
View dataset ↗
26 datapoints
SFT
Audio
Generalist
Voice SFT Samples
Scripted and unscripted multilingual speech recordings with verbatim transcripts for training voice-enabled speech and language AI systems.
View dataset ↗
11 datapoints
SFT
Audio
Generalist
Acoustic (Text-to-Audio) SFT
Indic scripted and spontaneous speech recordings across varied environments with word-for-word transcripts for speech generation and recognition models.
View dataset ↗
6 datapoints
Eval
STEM PhD
TEXT
Deep Research Agent Study - PCMB
Physics, chemistry, math, and biology prompts testing whether research agents propose novel next steps beyond existing literature.
View dataset ↗
5 datapoints
Eval
STEM PhD
TEXT
LongHorizon STEM Expert Tasks
Expert-steered multi-turn STEM problems with pre-registered expected directions for evaluating model collaboration under scope changes and injected constraints.
View dataset ↗
5 datapoints
SFT
Image
Visual Designers
GUI grounding based code generation
Broken-and-fixed app screenshots across finance, health, media, board management, and travel paired with golden React code for UI repair.
View dataset ↗
24 datapoints
SFT
Image
Generalist
Image-gen SFT
Input-output image pairs with editing prompts spanning multiple categories for training image-to-image editing models.
View dataset ↗
20 datapoints
SFT
Image
Generalist
Image-gen Thinking traces
Reasoning traces paired with prompts and resulting images, human-validated for training image models to reason before generating.
View dataset ↗
15 datapoints
Eval
Video
Generalist
Multi-image reference to Video generation
Multi-reference images and prompts paired with 10-second generated videos scored across six dimensions for video-generation evaluation.
View dataset ↗
30 datapoints
Eval
Image
Generalist
Spatial Counterfactual occlusion and rearrangement evaluation (SCORE)
Six spatial-editing workflows, including transparency, reflection, and object swapping, RLVR-scored for evaluating model spatial reasoning accuracy.
View dataset ↗
35 datapoints
SFT
Video
Generalist
Video-to-video SFT Dataset
50,000 source-target video pairs with editing instructions spanning object changes, backgrounds, style, and effects for video-editing models.
View dataset ↗
30 datapoints
Eval
Video
Generalist
Visual-Captioning
AI-generated captions paired with human corrections across video categories including industrial work, sports, and tutorials for caption-accuracy benchmarking.
View dataset ↗
11 datapoints
SFT
Audio
Generalist
Acoustic (Text-to-Audio) SFT
Indic scripted and spontaneous speech recordings across varied environments with word-for-word transcripts for speech generation and recognition models.
View dataset ↗
8 datapoints
Eval
Image
Generalist
Rubrics as RL
Labeled taxonomy of eighteen rubric-writing skills across five categories and three proficiency levels for competency mapping and annotator calibration.
View dataset ↗
15 datapoints
eval
video
Generalist
Multi-image reference to Video generation
Multi-reference images and edit prompts paired with human-scored 10-second videos assessed across six dimensions for video-generation evaluation.
View dataset ↗
The Deccan AI Standard for Datasets
Explore datasets designed to be challenging, trustworthy, legally usable, and ready to scale.
Rigorous, Multi-Step Quality Control
Every sample passes through a purpose-built QC pipeline: human-led or hybrid approach using custom LLM judges with expert human validation.
Model-Breaking by Design
Datasets are benchmarked against five frontier models and accepted only when they expose meaningful failures in at least three.
Clean, Verifiable Usage Rights
Privately acquired data is sourced with clearly documented provenance and usage rights.
Built to Scale
Every creation pipeline is engineered to expand with your volume, complexity, and evolving model requirements.



.webp)
