Skip to content
15 multimodal datasets to know in 2026
Published: March 23, 2026 Last update: August 14, 2026

15 multimodal datasets to know in 2026

Devika Garg
Devika Garg

Multimodal datasets combine multiple data types like text, images, audio, video, genomics, and clinical data to create richer AI training resources. These datasets are crucial for multimodal deep learning, which requires integrating multiple data sources to enhance performance in tasks such as image captioning, medical diagnostics, scientific discovery, autonomous systems, and cross-modal analysis.

Professionals working with multimodal data need access to high-quality datasets that provide the comprehensive information required for training sophisticated AI models. The datasets featured in this guide represent the most valuable resources available for advancing multimodal machine learning research and applications across computer vision, natural language processing, and life sciences.

What is a multimodal dataset?

A multimodal dataset contains information from multiple data modalities such as text, images, audio, video, genomics, clinical records, and sensor data. Multimodal datasets bring different modalities together, helping researchers build AI models that can understand complex environments and reason across multiple information sources.

Multimodal datasets are the digital equivalent of our senses. Just as we use sight, sound, and touch to interpret the world, these datasets combine various data formats to offer a richer understanding of content and enable more sophisticated analysis.

Multimodal datasets enable AI systems to process information more holistically by combining complementary data types. This approach mirrors human cognition, where we naturally integrate information from multiple sources to understand our environment.

Dataset Size License Authors Ideal for
TCGA20K+ samples, 2.5 PBOpen accessNCI/NHGRICancer genomics
CLEVR1M images + questionsBSDJohnson et al.Visual reasoning
WIT37.6M image-text pairsCC BY-SA 3.0Srinivasan et al.Multilingual learning
LAION-5B5.85B image-text pairsVariousLAION collectiveLarge-scale training
MS COCO330K imagesCC BY 4.0MicrosoftObject detection
VQA265K images + questionsCC BY 4.0Antol et al.Question answering
MIMIC-IV400K+ admissionsPhysioNet LicenseJohnson et al.Healthcare AI
UK Biobank500K+ participantsApproved accessUK BiobankPopulation health
Human Cell AtlasMillions of cellsOpen accessHCA ConsortiumSingle-cell biology
TCIA30M+ imagesVariousClark et al.Medical imaging
MLOmics8,314 cancer samplesOpen accessChen et al.Cancer ML
Kinetics-700650K video clipsYouTube ToSCarreira et al.Action recognition
Flickr30k31K images + captionsAcademic useYoung et al.Image-text retrieval
ENCODE1,600+ experimentsOpen accessENCODE ConsortiumGene regulation
Allen Brain AtlasMulti-brain regionsOpen accessAllen InstituteNeuroscience

To analyze multimodal datasets, specialized platforms and tools are required that can handle the complexity of multiple data types simultaneously. Tile.ai makes multimodal data AI-ready in place: each source becomes a Tile that carries the operations its own format needs, so genomics, imaging and clinical records are queried where they already live, in any format, with no pipelines and no copies.

1) TCGA (The Cancer Genome Atlas)

TCGA is a landmark cancer genomics program that molecularly characterized over 20,000 primary cancers and matched normal samples spanning 33 cancer types. This joint effort between NCI and the National Human Genome Research Institute generated over 2.5 petabytes of genomic, epigenomic, transcriptomic, and proteomic data.

TCGA combines multiple omics modalities including whole genome sequencing, RNA sequencing, DNA methylation analysis, protein expression, and comprehensive clinical data. The dataset has fundamentally transformed cancer research by enabling integrated analysis across molecular data types.

Size20,000+ samples across 33 cancer types, 2.5 petabytes of data
LicenseOpen access through dbGaP and GDC
Access the datasetGenomic Data Commons
Ideal forCancer research, multi-omics integration, precision medicine, biomarker discovery, and therapeutic target identification.
Grid of TCGA tumour types compared across molecular platforms.
Figure 1: Integrated data set for comparing and contrasting multiple tumor types. (https://www.nature.com/articles/ng.2764)
GDC data portal summary showing case counts by primary tumour site.
Figure 2: Cases by major primary site - Data portal summary. (https://portal.gdc.cancer.gov/)

2) CLEVR (Compositional Language and Elementary Visual Reasoning)

CLEVR is a multimodal dataset designed to evaluate a machine learning model's ability to reason about the physical world using both visual information and natural language. It is a synthetic multimodal dataset created to test AI systems' ability to perform complex reasoning about visual scenes.

CLEVR combines visual and textual modalities through rendered 3D scenes containing various objects with distinct properties like shape, size, color, and material. The textual component consists of questions posed in natural language about these scenes, challenging models to understand relationships and properties.

Size1 million images with corresponding questions
LicenseBSD
Access the datasetStanford CLEVR Dataset
Ideal forResearchers developing visual reasoning systems, computer vision applications requiring spatial understanding, and AI models that need to process complex visual-linguistic relationships.
CLEVR field guide: object shapes and attributes, example questions with their functional programs, and the function catalog.
Figure 3: A field guide to the CLEVR universe. Left: Shapes, attributes, and spatial relationships. Center: Examples of questions and their associated functional programs. Right: Function catalog for building questions. (https://arxiv.org/pdf/1612.06890)

3) WIT (Wikipedia-based Image Text Dataset)

WIT primarily focuses on tasks involving the relationship between images and their textual descriptions. Some key applications are Image-Text Retrieval to retrieve images using text query, Image Captioning to generate captions for unseen images, and Multilingual Learning that can understand and connect images to text descriptions in various languages.

WIT provides a massive collection of image-text pairs extracted from Wikipedia articles across 108 languages, making it invaluable for multilingual multimodal applications.

Size37.6 million entity-rich image-text examples with 11.5 million unique images
LicenseCreative Commons Attribution-ShareAlike 3.0 Unported
Ideal forMultilingual AI applications, cross-modal retrieval systems, image captioning models, and researchers working on Wikipedia-scale knowledge representation.
A Wikipedia image from WIT shown with its full set of text annotations.
Figure 4: WIT image-text example with all text annotations. (https://arxiv.org/pdf/2103.01913)

4) LAION-5B

Developed by the LAION non-profit collective, LAION-5B is one of the largest open multimodal datasets available. It contains 5.85 billion CLIP-filtered image-text pairs extracted from the web, with 2.32 billion English pairs and the rest spanning multiple languages.

LAION-5B democratizes access to large-scale data used for training foundation models like CLIP and Stable Diffusion, providing researchers with web-scale multimodal data.

Size5.85 billion image-text pairs
LicenseVarious (depends on source)
Access the datasetLAION Dataset
Ideal forTraining large-scale foundation models, text-to-image generation, multimodal representation learning, and researchers requiring massive datasets for model development.
Sample image-caption pairs retrieved from LAION-5B by CLIP nearest-neighbour search.
Figure 5: LAION-5B examples. Sample images from a nearest neighbor search in LAION-5B using CLIP embeddings. The image and caption (C) are the first results for the query (Q). (https://arxiv.org/pdf/2210.08402)

5) MS COCO (Microsoft Common Objects in Context)

MS COCO provides comprehensive annotations for object detection, segmentation, and captioning tasks. The dataset contains images with multiple objects in natural contexts, accompanied by detailed annotations and captions.

Size330,000 images with over 2.5 million labeled instances
LicenseCreative Commons Attribution 4.0
Access the datasetCOCO Dataset
Ideal forObject detection research, image segmentation, caption generation, and computer vision benchmarking.
Comparison of object-recognition tasks, ending with MS COCO's instance segmentation of everyday scenes.
Figure 6: Previous object recognition datasets focus on (a), (b) and (c), MS COCO focuses on (d) segmenting individual object instances, by introducing a large, richly-annotated dataset comprised of images depicting complex everyday scenes of common objects in their natural context. (https://arxiv.org/pdf/1405.0312)

6) VQA (Visual Question Answering)

VQA combines visual and textual information to enable AI systems to answer questions about images. The dataset requires models to understand both visual content and natural language questions.

Size265,016 images with over 614,000 questions
LicenseCreative Commons Attribution 4.0
Access the datasetVQA Dataset
Ideal forVisual question answering research, multimodal reasoning systems, and educational AI applications.
Example open-ended VQA questions about photographs, each needing commonsense as well as visual understanding.
Figure 7: Examples of free-form, open-ended questions collected for images via Amazon Mechanical Turk. Note that commonsense knowledge is needed along with a visual understanding of the scene to answer many questions. (https://arxiv.org/pdf/1505.00468)

7) MIMIC-IV

MIMIC-IV provides multimodal medical data including clinical notes, lab results, vital signs, medications, and procedures from intensive care unit patients. The dataset enables comprehensive healthcare AI research by combining structured and unstructured clinical data.

MIMIC-IV represents one of the most complete multimodal clinical datasets available, integrating time-series physiological data with rich textual clinical documentation and structured medical records.

Size400,000+ hospital admissions with comprehensive clinical data
LicensePhysioNet Credentialed Health Data License
Access the datasetMIMIC-IV PhysioNet
Ideal forHealthcare AI research, clinical decision support systems, medical predictive modeling, and natural language processing in healthcare.
Flow diagram of how the MIMIC-IV dataset is assembled from hospital source systems.
Figure 8: An overview of the development process for MIMIC-IV. (https://www.nature.com/articles/s41597-022-01899-x.pdf)

8) UK Biobank

UK Biobank provides genomic data, medical imaging (MRI, retinal), proteomics, clinical records, and lifestyle data from 500,000+ participants with longitudinal follow-up. This massive prospective cohort study combines genetic, environmental, and lifestyle factors with detailed health outcomes.

The dataset includes whole genome sequencing, multi-organ MRI imaging, retinal photography, and comprehensive phenotyping data, making it invaluable for understanding gene-environment interactions in health and disease.

Size500,000+ participants with genomic, imaging, and clinical data
LicenseApproved researcher access
Access the datasetUK Biobank
Ideal forPopulation genomics, precision medicine, epidemiological studies, and gene-environment interaction research.
Summary of the UK Biobank resource and the content of its genotyping array.
Figure 9: Summary of the UK Biobank resource and genotyping array content. (http://nature.com/articles/s41586-018-0579-z)

9) Human Cell Atlas (HCA)

The Human Cell Atlas integrates single-cell RNA sequencing, spatial transcriptomics, proteomics, and imaging data to create comprehensive maps of all human cells. This international collaboration aims to map every cell type in the human body across development, health, and disease.

HCA combines cutting-edge single-cell technologies with spatial information to understand cellular diversity, tissue organization, and developmental processes at unprecedented resolution.

Research paperThe Human Cell Atlas
SizeMillions of cells across multiple tissues and developmental stages
LicenseOpen access with data use agreements
Ideal forSingle-cell biology, developmental biology, disease research, and understanding cellular diversity in human tissues.
Human Cell Atlas illustration of cell types arranged by tissue structure.
Figure 10: Anatomy: cell types and tissue structure. (https://elifesciences.org/articles/27041)

10) TCIA (The Cancer Imaging Archive)

The compiled datasets encompass a broad spectrum of data modalities, such as radiology images (CT, MRI, PET), pathology slides, genomic data, and clinical records. This multimodal nature enables the integration of different data types to capture the intricacies of cancer.

TCIA provides medical imaging data paired with clinical information for cancer research and diagnosis development.

Size30+ million images across multiple cancer types
LicenseVarious (most are open access)
Access the datasetThe Cancer Imaging Archive
Ideal forMedical AI research, cancer diagnosis systems, radiological analysis, and imaging-genomics integration.

11) MLOmics

MLOmics is an open cancer multi-omics database specifically designed for machine learning applications. The dataset contains genomic, transcriptomic, proteomic, and clinical data from cancer patients, with extensive preprocessing and baseline models for immediate use in ML workflows.

MLOmics addresses the gap between raw omics data and machine learning-ready datasets by providing standardized, well-annotated multimodal cancer data with comprehensive baseline evaluations.

Size8,314 patient samples covering 32 cancer types with four omics types
LicenseOpen access
Access the datasetMLOmics Database
Ideal forCancer machine learning research, multi-omics integration, biomarker discovery, and developing AI models for precision oncology.
Workflow for building MLOmics from raw cancer multi-omics data.
Figure 11: Schematic workflow of creating the MLOmics. (https://www.nature.com/articles/s41597-025-05235-x.pdf)
Structure of the MLOmics resource: omics types, sample groups and baseline models.
Figure 12: Schematic of MLOmics resources structure. (https://www.nature.com/articles/s41597-025-05235-x.pdf)

12) Kinetics-700

Kinetics-700 contains human action videos across 700 action categories, providing rich temporal visual data for action recognition and video understanding research.

Size650,000 video clips
LicenseYouTube terms of service
Access the datasetKinetics dataset (CVDF)
Ideal forAction recognition research, video classification, and temporal modeling in computer vision.
Example action classes from Kinetics-700, several of which need more than one frame to identify.
Figure 13: Example classes from the Kinetics dataset. Note that in some cases a single image is not enough for recognizing the action or distinguishing classes. (https://arxiv.org/pdf/1705.06950)

13) Flickr30k

Flickr30k contains images from Flickr paired with human-written captions, providing high-quality image-text associations for multimodal learning.

Size31,000 images with 158,000 captions
LicenseAcademic research
Access the datasetFlickr30k Dataset
Ideal forImage-text retrieval, caption generation, and visual-semantic understanding research.
Two Flickr30k photographs, each shown with its five human-written captions.
Figure 14: Two images from the data set and their five captions. (https://aclanthology.org/Q14-1006.pdf)

14) ENCODE

ENCODE provides comprehensive data on gene regulation through ChIP-seq, RNA-seq, chromatin accessibility assays, and histone modification mapping across cell types and conditions. The project aims to identify all functional elements in the human genome.

ENCODE integrates multiple genomic assays to create comprehensive maps of gene regulatory networks, combining transcriptional, epigenetic, and chromatin accessibility data.

Size1,600+ experiments across multiple cell types and conditions
LicenseOpen access
Access the datasetENCODE Portal
Ideal forGene regulation research, epigenomics studies, regulatory network analysis, and understanding functional genomics.
Genome-wide segmentation integrating ENCODE assay tracks.
Figure 15: Integration of ENCODE data by genome-wide segmentation. (https://www.nature.com/articles/nature11247)

15) Allen Brain Atlas

The Allen Brain Atlas combines gene expression data, neuroimaging, connectivity mapping, and cell type characterization across brain regions and developmental stages. This comprehensive resource provides multimodal data for understanding brain structure and function.

The atlas integrates anatomical, molecular, and functional data to create detailed maps of brain organization across species and developmental time points.

SizeMultiple brain regions across development and disease states
LicenseOpen access
Access the datasetAllen Brain Atlas Portal
Ideal forNeuroscience research, brain development studies, neurological disease research, and systems neuroscience.
Adult mouse brain, 3D coronal view from the Allen Brain Atlas.
Figure 16: Adult mouse 3D coronal.

Are there free multimodal datasets?

Yes, there are numerous free multimodal datasets available for research and development purposes. Many prominent datasets in this guide are freely accessible, including:

  • TCGA - Available through open access portals
  • CLEVR - Released under BSD license
  • Human Cell Atlas - Open access with data use agreements

Free datasets may have limitations such as restricted commercial use, academic-only licensing, or requirements for attribution. Researchers can find free datasets through academic institutions, government organizations, and open-source communities.

How can you find multimodal datasets?

You can find multimodal datasets through specialized repositories, academic institutions, and research organizations.

Academic repositories

Papers with Code, Google Dataset Search, and university research labs

Government sources

NIH data repositories, cancer genomics portals, and national research initiatives

Industry platforms

Kaggle, Hugging Face, and company research divisions

Specialized collections

Medical imaging archives, genomics databases, and domain-specific repositories

How can you analyze multimodal datasets?

To analyze multimodal datasets, you need specialized platforms and tools capable of handling heterogeneous data types simultaneously. Traditional database systems struggle with the complexity and scale of multimodal data, particularly in life sciences where datasets may include genomics, imaging, and clinical data.

Modern multimodal analysis requires platforms that can efficiently store, index, and query across different data modalities. These systems must handle the temporal synchronization between modalities and provide consistent access across disparate data types.

Tile.ai provides a comprehensive platform for multimodal data analysis as the Enterprise AI Data Substrate, a governed layer that sits alongside the data and applications you already use. The platform enables organizations to discover, query, and analyze complex multimodal datasets efficiently where they already live, supporting everything from medical imaging combined with genomic data to single-cell multiomics with spatial information.

For organizations looking to implement multimodal data analysis capabilities, Tile.ai works alongside the platforms already in place rather than replacing them. Request access to see your own data queried where it sits, or read the multimodal data platform buyer's guide if you are still comparing options.

Which industries use multimodal datasets?

Multimodal datasets are transforming industries by enabling more comprehensive analysis and decision-making across sectors.

  • Academic research: Universities combine imaging, genomics, and clinical data for breakthrough medical discoveries
  • Biotechnology: Single-cell multiomics integrates genomics, proteomics, and spatial data for understanding cellular biology
  • Diagnostics: Medical device manufacturers combine imaging data with AI models for automated diagnosis
  • Healthcare: Multimodal datasets can combine medical imaging, genomic data, patient records, and clinical notes for improved diagnosis and treatment planning. This multimodal approach enables the integration of different data types to capture the intricacies of cancer and other complex medical conditions.
  • Healthcare systems: Hospitals integrate electronic health records, medical imaging, and laboratory data for clinical decision support
  • Pharmaceutical: Drug discovery combines molecular data, imaging, genomics, and clinical trial data for precision medicine approaches

Additional industries adopting multimodal approaches include agriculture (combining satellite imagery with genomics data for crop improvement), environmental science (integrating sensor data with imaging for climate research), and materials science (combining microscopy with molecular data for material discovery).

Multimodal datasets enable these industries to build AI systems that mirror human-like understanding by processing information from multiple sources simultaneously. Organizations implementing multimodal approaches often see significant improvements in accuracy and insights compared to single-modality systems.

Ready to put multimodal data to work in your organization? Tile.ai makes multi-omics, research, clinical and manufacturing data AI-ready — in place, in the systems it already lives in. Request access: point us to a data source and a service, and we'll have your agents querying them in place.

Show me my data, activated.

Point us to a data source and a service, and we'll have your agents querying them in place — in an afternoon, not a quarter.