unified_qa  |  TensorFlow Datasets Skip to main content Install Learn Introduction New to TensorFlow? Tutorials Learn how to use TensorFlow with end-to-end examples Guide Learn framework concepts and components Learn ML Educational resources to master your path with TensorFlow API TensorFlow (v2.16.1) Versions… TensorFlow.js TensorFlow Lite TFX Resources LIBRARIES TensorFlow.js Develop web ML applications in JavaScript TensorFlow Lite Deploy ML on mobile, microcontrollers and other edge devices TFX Build production ML pipelines All libraries Create advanced models and extend TensorFlow RESOURCES Models & datasets Pre-trained models and datasets built by Google and the community Tools Tools to support and accelerate TensorFlow workflows Responsible AI Resources for every stage of the ML workflow Recommendation systems Build recommendation systems with open source tools Community Groups User groups, interest groups and mailing lists Contribute Guide for contributing to code and documentation Blog Stay up to date with all things TensorFlow Forum Discussion platform for the TensorFlow community Why TensorFlow About Case studies / English Español – América Latina Français Indonesia Italiano Polski Português – Brasil Tiếng Việt Türkçe Русский עברית العربيّة فارسی हिंदी বাংলা ภาษาไทย 中文 – 简体 日本語 한국어 GitHub Sign in Datasets Overview Catalog Community Catalog Guide API Install Learn More API More Resources More Overview Catalog Community Catalog Guide API Community More Why TensorFlow More GitHub Overview Dataset Collections longt5 xtreme 3d aflw2k3d smallnorb smartwatch_gestures Abstractive text summarization aeslc billsum booksum (manual) multi_news newsroom (manual) reddit reddit_tifu samsum (manual) scientific_papers Age dices wake_vision Anomaly detection ag_news_subset caltech101 kddcup99 lost_and_found stl10 Audio accentdb common_voice crema_d dementiabank (manual) fuss groove gtzan gtzan_music_speech librispeech libritts ljspeech nsynth savee (manual) speech_commands spoken_digit tedlium user_libri_audio vctk voxceleb (manual) voxforge (manual) xtreme_s yes_no Biology ai2dcaption ogbg_molpcba Categorical dices sift1m wake_vision Common sense reasoning ai2_arc_with_ir arc covr natural_questions openbookqa Computer science robomimic_mg robomimic_mh robomimic_ph smart_buildings Conditional image generation imagenet2012 (manual) imagenet2012_subset (manual) webvid (manual) Coreference resolution clevr D4rl d4rl_adroit_door d4rl_adroit_hammer d4rl_adroit_pen d4rl_adroit_relocate d4rl_antmaze d4rl_mujoco_ant d4rl_mujoco_halfcheetah d4rl_mujoco_hopper d4rl_mujoco_walker2d Density estimation caltech101 celeb_a_hq (manual) imagenet2012 (manual) imagenet2012_subset (manual) Dependency parsing universal_dependencies xtreme_pos Dialog act labeling bot_adversarial_dialogue Dialogue bot_adversarial_dialogue databricks_dolly dices Document summarization aeslc booksum (manual) newsroom (manual) scientific_papers Facial attributes wake_vision Fine grained image classification caltech101 oxford_flowers102 oxford_iiit_pet stanford_dogs stl10 sun397 wake_vision Gender dices wake_vision Graph ogbg_molpcba reddit Graphs cardiotox Health pneumonia_mnist Image abstract_reasoning (manual) aflw2k3d ai2dcaption bccd beans bee_dataset bigearthnet binarized_mnist binary_alpha_digits caltech101 celeb_a celeb_a_hq (manual) cityscapes (manual) clevr clic coil100 covr div2k downsampled_imagenet dsprites flic imagenet2012 (manual) imagenet2012_corrupted (manual) imagenet2012_fewshot (manual) imagenet2012_multilabel (manual) imagenet2012_real (manual) imagenet2012_subset (manual) imagenet_a imagenet_lt (manual) imagenet_pi (manual) imagenet_r imagenet_resized imagenet_sketch imagenet_v2 imagenette imagewang kitti lfw lost_and_found lsun lvis malaria nyu_depth_v2 open_images_challenge2019_detection open_images_v4 oxford_flowers102 oxford_iiit_pet pass patch_camelyon pet_finder places365_small placesfull plant_leaves plant_village plantae_k pneumonia_mnist quickdraw_bitmap ref_coco (manual) resisc45 (manual) robomimic_mg robomimic_mh robomimic_ph rock_paper_scissors s3o4d scene_parse150 shapes3d siscore smallnorb so2sat stanford_dogs stanford_online_products stl10 sun397 svhn_cropped symmetric_solids tf_flowers the300w_lp wake_vision Image classification abstract_reasoning (manual) bigearthnet caltech101 caltech_birds2010 caltech_birds2011 cars196 cassava cats_vs_dogs celeb_a chexpert (manual) cifar10 cifar100 cifar100_n (manual) cifar10_1 cifar10_corrupted cifar10_h cifar10_n (manual) citrus_leaves cmaterdb colorectal_histology colorectal_histology_large controlled_noisy_web_labels (manual) curated_breast_imaging_ddsm (manual) cycle_gan deep_weeds diabetic_retinopathy_detection (manual) dmlab domainnet dtd emnist eurosat fashion_mnist flic food101 geirhos_conflict_stimuli horses_or_humans i_naturalist2017 i_naturalist2018 i_naturalist2021 imagenet2012 (manual) imagenet2012_subset (manual) imagenet_resized imagenet_sketch imagenette imagewang kitti kmnist lvis malaria mnist mnist_corrupted omniglot open_images_v4 oxford_flowers102 oxford_iiit_pet places365_small plant_village pneumonia_mnist resisc45 (manual) siscore smallnorb stanford_dogs stanford_online_products stl10 sun397 svhn_cropped uc_merced visual_domain_decathlon wake_vision Image clustering imagenet2012 (manual) imagenet2012_subset (manual) stanford_dogs stl10 Image compression imagenet2012 (manual) imagenet2012_subset (manual) imagenet_resized oxford_iiit_pet patch_camelyon stl10 Image generation binarized_mnist celeb_a celeb_a_hq (manual) cityscapes (manual) clevr imagenet2012 (manual) imagenet2012_subset (manual) oxford_flowers102 stanford_dogs stl10 Image segmentation segment_anything (manual) Image super resolution celeb_a_hq (manual) div2k kitti Image to image translation celeb_a_hq (manual) cityscapes (manual) kitti scene_parse150 Instance segmentation cityscapes (manual) lvis nyu_depth_v2 segment_anything (manual) Language modeling big_patent billsum blimp databricks_dolly dices dolma e2e_cleaned irc_disentanglement lambada librispeech_lm lm1b math_qa opus paws_wiki paws_x_wiki pg19 reddit samsum (manual) snli squad tedlium Linguistic acceptability bot_adversarial_dialogue Machine translation mlqa opus Monolingual ag_news_subset ai2_arc_with_ir arc beir booksum (manual) bool_q bot_adversarial_dialogue covr dices dolma e2e_cleaned imdb_reviews kitti lambada librispeech librispeech_lm libritts ljspeech lm1b natural_questions natural_questions_open openbookqa paws_wiki plant_village quac race real_toxicity_prompts reddit savee (manual) schema_guided_dialogue sci_tail scicite scientific_papers sentiment140 snli speech_commands spoken_digit squad story_cloze (manual) tedlium trec trivia_qa Movies and tv shows opinion_abstracts Multilingual librispeech mlqa paws_x_wiki Natural language inference anli covr dices paws_wiki sci_tail snli Natural language understanding ag_news_subset anli beir bool_q clevr covr databricks_dolly dices imdb_reviews math_dataset math_qa mlqa multi_news natural_instructions natural_questions natural_questions_open openbookqa opus paws_wiki paws_x_wiki pg19 piqa qasc quac race schema_guided_dialogue sci_tail sentiment140 snli squad story_cloze (manual) trec trivia_qa Nearest neighbors deep1b glove100_angular News multi_news Object detection coco coco_captions covr flic kitti lvis open_images_v4 voc waymo_open_dataset wider_face Open domain question answering databricks_dolly natural_questions squad trivia_qa Out of distribution detection stl10 Question answering beir bool_q clevr coqa cosmos_qa databricks_dolly math_dataset math_qa mlqa natural_instructions natural_questions natural_questions_open openbookqa piqa qasc quac race squad story_cloze (manual) trivia_qa tydi_qa web_questions xquad Question generation natural_questions trivia_qa Race national ethnic origin dices Ranking istella mslr_web yahoo_ltrc (manual) Reading comprehension opus pg19 qasc quac race squad trivia_qa Recommendation criteo hillstrom Reinforcement learning robomimic_mg robomimic_mh robomimic_ph smart_buildings Rgb d nyu_depth_v2 Rl unplugged rlu_atari rlu_atari_checkpoints rlu_atari_checkpoints_ordered rlu_control_suite rlu_dmlab_explore_object_rewards_few rlu_dmlab_explore_object_rewards_many rlu_dmlab_rooms_select_nonmatching_object rlu_dmlab_rooms_watermaze rlu_dmlab_seekavoid_arena01 rlu_locomotion rlu_rwrl Rlds locomotion robosuite_panda_pick_place_can Robotics aloha_mobile asimov_dilemmas_auto_val asimov_dilemmas_scifi_train asimov_dilemmas_scifi_val asimov_injury_val asimov_multimodal_auto_val asimov_multimodal_manual_val asimov_v2_constraints_with_rationale asimov_v2_constraints_without_rationale asimov_v2_injuries asimov_v2_videos asu_table_top_converted_externally_to_rlds austin_buds_dataset_converted_externally_to_rlds austin_sailor_dataset_converted_externally_to_rlds austin_sirius_dataset_converted_externally_to_rlds bc_z berkeley_autolab_ur5 berkeley_cable_routing berkeley_fanuc_manipulation berkeley_gnm_cory_hall berkeley_gnm_recon berkeley_gnm_sac_son berkeley_mvp_converted_externally_to_rlds berkeley_rpt_converted_externally_to_rlds bridge bridge_data_msr cmu_franka_exploration_dataset_converted_externally_to_rlds cmu_play_fusion cmu_stretch columbia_cairlab_pusht_real conq_hose_manipulation dlr_edan_shared_control_converted_externally_to_rlds dlr_sara_grid_clamp_converted_externally_to_rlds dlr_sara_pour_converted_externally_to_rlds dobbe eth_agent_affordances fmb fractal20220817_data iamlab_cmu_pickup_insert_converted_externally_to_rlds imperialcollege_sawyer_wrist_cam io_ai_tech jaco_play kaist_nonprehensile_converted_externally_to_rlds kuka maniskill_dataset_converted_externally_to_rlds mimic_play mt_opt nyu_door_opening_surprising_effectiveness nyu_franka_play_dataset_converted_externally_to_rlds nyu_rot_dataset_converted_externally_to_rlds plex_robosuite robo_ai_u_r5e robo_set roboturk spoc_robot stanford_hydra_dataset_converted_externally_to_rlds stanford_kuka_multimodal_dataset_converted_externally_to_rlds stanford_mask_vit_converted_externally_to_rlds stanford_robocook_converted_externally_to_rlds taco_play tidybot tokyo_u_lsmo_converted_externally_to_rlds toto ucsd_kitchen_dataset_converted_externally_to_rlds ucsd_pick_and_place_dataset_converted_externally_to_rlds uiuc_d3field usc_cloth_sim_converted_externally_to_rlds utaustin_mutex utokyo_pr2_opening_fridge_converted_externally_to_rlds utokyo_pr2_tabletop_manipulation_converted_externally_to_rlds utokyo_saytap_converted_externally_to_rlds utokyo_xarm_bimanual_converted_externally_to_rlds utokyo_xarm_pick_and_place_converted_externally_to_rlds vima_converted_externally_to_rlds viola Scene classification bigearthnet places365_small resisc45 (manual) Semantic segmentation bigearthnet cityscapes (manual) kitti lost_and_found nyu_depth_v2 open_images_v4 places365_small ref_coco (manual) scene_parse150 segment_anything (manual) so2sat Sentiment analysis imdb_reviews sentiment140 Sequence modeling databricks_dolly smart_buildings Sequence to sequence language modeling big_patent billsum databricks_dolly math_qa opus paws_wiki reddit samsum (manual) snli Speech librispeech libritts speech_commands Speech recognition accentdb librispeech speech_commands tedlium Structured cherry_blossoms covid19 cs_restaurants dart diamonds forest_fires genomics_ood german_credit_numeric higgs howell iris movie_lens movielens web_graph web_nlg wiki_bio wiki_table_questions wiki_table_text wine_quality Summarization cnn_dailymail covid19sum (manual) gigaword gov_report wikihow (manual) xsum (manual) Table to text generation e2e_cleaned Tabular ble_wind_field efron_morris75 kddcup99 opinion_abstracts radon simpte (manual) titanic Text abstract_reasoning (manual) aeslc ag_news_subset ai2_arc ai2_arc_with_ir amazon_us_reviews anli answer_equivalence arc asqa asset assin2 bccd beir big_patent billsum blimp booksum (manual) bool_q bot_adversarial_dialogue bucc c4 (manual) c4_wsrs caltech101 cfq cityscapes (manual) civil_comments clevr clinc_oos conll2002 conll2003 corr2cause cos_e covr databricks_dolly definite_pronoun_resolution dices doc_nli dolma dolphin_number_word drop dsprites e2e_cleaned eraser_multi_rc esnli flic gap gem glue goemotions gpt3 gsm8k hellaswag imdb_reviews irc_disentanglement kitti lambada lfw librispeech librispeech_lm libritts ljspeech lm1b lost_and_found math_dataset math_qa mctaco media_sum (manual) mlqa movie_rationales mrqa multi_news multi_nli multi_nli_mismatch natural_instructions natural_questions natural_questions_open newsroom (manual) open_images_challenge2019_detection open_images_v4 openbookqa opinion_abstracts opinosis opus oxford_flowers102 oxford_iiit_pet para_crawl patch_camelyon paws_wiki paws_x_wiki penguins pet_finder pg19 piqa places365_small placesfull plant_leaves plant_village plantae_k protein_net q_re_cc qa4mre qasc quac quality race real_toxicity_prompts reddit_disentanglement (manual) reddit_tifu ref_coco (manual) resisc45 (manual) robonet rock_you salient_span_wikipedia samsum (manual) scan schema_guided_dialogue sci_tail scicite scientific_papers scrolls sentiment140 smallnorb snli spoken_digit squad squad_question_generation stanford_dogs star_cfq story_cloze (manual) summscreen sun397 super_glue svhn_cropped tatoeba ted_hrlr_translate ted_multi_translate tedlium tiny_shakespeare trec trivia_qa unified_qa universal_dependencies unnatural_instructions user_libri_text webvid (manual) wiki40b wiki_dialog wikiann wikipedia wikipedia_toxicity_subtypes winogrande wordnet wsc273 xnli xtreme_pawsx xtreme_xnli yelp_polarity_reviews Text classification ag_news_subset bool_q bot_adversarial_dialogue dices imdb_reviews natural_instructions paws_wiki paws_x_wiki sentiment140 trec Text classification toxicity prediction bot_adversarial_dialogue dices real_toxicity_prompts Text generation aeslc big_patent billsum booksum (manual) bool_q databricks_dolly e2e_cleaned lambada lm1b math_qa mctaco natural_instructions natural_questions newsroom (manual) openbookqa piqa race real_toxicity_prompts reddit reddit_tifu samsum (manual) schema_guided_dialogue scientific_papers squad story_cloze (manual) trivia_qa Text simplification wiki_auto Text summarization aeslc big_patent billsum booksum (manual) databricks_dolly multi_news newsroom (manual) reddit reddit_tifu samsum (manual) scientific_papers Time series robomimic_mg robomimic_mh robomimic_ph smart_buildings smartwatch_gestures Token classification universal_dependencies xtreme_pos Tracking smartwatch_gestures Trajectory robomimic_mg robomimic_mh robomimic_ph Translate flores mtnt wmt13_translate (manual) wmt14_translate (manual) wmt15_translate (manual) wmt16_translate (manual) wmt17_translate (manual) wmt18_translate (manual) wmt19_translate (manual) wmt_t2t_translate (manual) Uncategorized duke_ultrasound lbpp qm9 Unsupervised anomaly detection caltech101 Video abstract_reasoning (manual) bair_robot_pushing_small davis flic moving_mnist robonet starcraft_video tao (manual) ucf101 webvid (manual) youtube_vis (manual) Vision language gref (manual) grounded_scan laion400m (manual) wit wit_kaggle (manual) Introduction Tutorials Guide Learn ML TensorFlow (v2.16.1) Versions… TensorFlow.js TensorFlow Lite TFX LIBRARIES TensorFlow.js TensorFlow Lite TFX All libraries RESOURCES Models & datasets Tools Responsible AI Recommendation systems Groups Contribute Blog Forum About Case studies TFDS now supports the Croissant 🥐 format! Read the documentation to know more. TensorFlow Resources Datasets Catalog unified_qa Stay organized with collections Save and categorize content based on your preferences. Description: The UnifiedQA benchmark consists of 20 main question answering (QA) datasets (each may have multiple versions) that target different formats as well as various complex linguistic phenomena. These datasets are grouped into several formats/categories, including: extractive QA, abstractive QA, multiple-choice QA, and yes/no QA. Additionally, contrast sets are used for several datasets (denoted with "contrastsets"). These evaluation sets are expert-generated perturbations that deviate from the patterns common in the original dataset. For several datasets that do not come with evidence paragraphs, two variants are included: one where the datasets are used as-is and another that uses paragraphs fetched via an information retrieval system as additional evidence, indicated with "_ir" tags. More information can be found at: https://github.com/allenai/unifiedqa Homepage: https://github.com/allenai/unifiedqa Source code: tfds.text.unifiedqa.UnifiedQA Versions: 1.0.0 (default): Initial release. Feature structure: FeaturesDict({ 'input': string, 'output': string, }) Feature documentation: Feature Class Shape Dtype Description FeaturesDict input Tensor string output Tensor string Supervised keys (See as_supervised doc): None Figure (tfds.show_examples): Not supported. unified_qa/ai2_science_elementary (default config) Config description: The AI2 Science Questions dataset consists of questions used in student assessments in the United States across elementary and middle school grade levels. Each question is 4-way multiple choice format and may or may not include a diagram element. This set consists of questions used for elementary school grade levels. Download size: 345.59 KiB Dataset size: 390.02 KiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 542 'train' 623 'validation' 123 Examples (tfds.as_dataframe): Citation: http://data.allenai.org/ai2-science-questions @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/ai2_science_middle Config description: The AI2 Science Questions dataset consists of questions used in student assessments in the United States across elementary and middle school grade levels. Each question is 4-way multiple choice format and may or may not include a diagram element. This set consists of questions used for middle school grade levels. Download size: 428.41 KiB Dataset size: 477.40 KiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 679 'train' 605 'validation' 125 Examples (tfds.as_dataframe): Citation: http://data.allenai.org/ai2-science-questions @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/ambigqa Config description: AmbigQA is an open-domain question answering task which involves finding every plausible answer, and then rewriting the question for each one to resolve the ambiguity. Download size: 2.27 MiB Dataset size: 3.04 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 19,806 'validation' 5,674 Examples (tfds.as_dataframe): Citation: @inproceedings{min-etal-2020-ambigqa, title = "{A}mbig{QA}: Answering Ambiguous Open-domain Questions", author = "Min, Sewon and Michael, Julian and Hajishirzi, Hannaneh and Zettlemoyer, Luke", booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.emnlp-main.466", doi = "10.18653/v1/2020.emnlp-main.466", pages = "5783--5797", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_easy Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "easy" questions. Download size: 1.24 MiB Dataset size: 1.42 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 2,376 'train' 2,251 'validation' 570 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_easy_dev Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "easy" questions. Download size: 1.24 MiB Dataset size: 1.42 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 2,376 'train' 2,251 'validation' 570 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_easy_with_ir Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "easy" questions. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 7.00 MiB Dataset size: 7.17 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 2,376 'train' 2,251 'validation' 570 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_easy_with_ir_dev Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "easy" questions. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 7.00 MiB Dataset size: 7.17 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 2,376 'train' 2,251 'validation' 570 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_hard Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "hard" questions. Download size: 758.03 KiB Dataset size: 848.28 KiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,172 'train' 1,119 'validation' 299 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_hard_dev Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "hard" questions. Download size: 758.03 KiB Dataset size: 848.28 KiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,172 'train' 1,119 'validation' 299 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_hard_with_ir Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "hard" questions. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 3.53 MiB Dataset size: 3.62 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,172 'train' 1,119 'validation' 299 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/arc_hard_with_ir_dev Config description: This dataset consists of genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. This set consists of "hard" questions. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 3.53 MiB Dataset size: 3.62 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,172 'train' 1,119 'validation' 299 Examples (tfds.as_dataframe): Citation: @article{clark2018think, title={Think you have solved question answering? try arc, the ai2 reasoning challenge}, author={Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind}, journal={arXiv preprint arXiv:1803.05457}, year={2018} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/boolq Config description: BoolQ is a question answering dataset for yes/no questions. These questions are naturally occurring ---they are generated in unprompted and unconstrained settings. Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context. The text-pair classification setup is similar to existing natural language inference tasks. Download size: 7.77 MiB Dataset size: 8.20 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 9,427 'validation' 3,270 Examples (tfds.as_dataframe): Citation: @inproceedings{clark-etal-2019-boolq, title = "{B}ool{Q}: Exploring the Surprising Difficulty of Natural Yes/No Questions", author = "Clark, Christopher and Lee, Kenton and Chang, Ming-Wei and Kwiatkowski, Tom and Collins, Michael and Toutanova, Kristina", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)", month = jun, year = "2019", address = "Minneapolis, Minnesota", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N19-1300", doi = "10.18653/v1/N19-1300", pages = "2924--2936", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/boolq_np Config description: BoolQ is a question answering dataset for yes/no questions. These questions are naturally occurring ---they are generated in unprompted and unconstrained settings. Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context. The text-pair classification setup is similar to existing natural language inference tasks. This version adds natural perturbations to the original version. Download size: 10.80 MiB Dataset size: 11.40 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 9,727 'validation' 7,596 Examples (tfds.as_dataframe): Citation: @inproceedings{khashabi-etal-2020-bang, title = "More Bang for Your Buck: Natural Perturbation for Robust Question Answering", author = "Khashabi, Daniel and Khot, Tushar and Sabharwal, Ashish", booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.emnlp-main.12", doi = "10.18653/v1/2020.emnlp-main.12", pages = "163--170", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/commonsenseqa Config description: CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains questions with one correct answer and four distractor answers. Download size: 1.79 MiB Dataset size: 2.19 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,140 'train' 9,741 'validation' 1,221 Examples (tfds.as_dataframe): Citation: @inproceedings{talmor-etal-2019-commonsenseqa, title = "{C}ommonsense{QA}: A Question Answering Challenge Targeting Commonsense Knowledge", author = "Talmor, Alon and Herzig, Jonathan and Lourie, Nicholas and Berant, Jonathan", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)", month = jun, year = "2019", address = "Minneapolis, Minnesota", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N19-1421", doi = "10.18653/v1/N19-1421", pages = "4149--4158", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/commonsenseqa_test Config description: CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains questions with one correct answer and four distractor answers. Download size: 1.79 MiB Dataset size: 2.19 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,140 'train' 9,741 'validation' 1,221 Examples (tfds.as_dataframe): Citation: @inproceedings{talmor-etal-2019-commonsenseqa, title = "{C}ommonsense{QA}: A Question Answering Challenge Targeting Commonsense Knowledge", author = "Talmor, Alon and Herzig, Jonathan and Lourie, Nicholas and Berant, Jonathan", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)", month = jun, year = "2019", address = "Minneapolis, Minnesota", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N19-1421", doi = "10.18653/v1/N19-1421", pages = "4149--4158", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/contrast_sets_boolq Config description: BoolQ is a question answering dataset for yes/no questions. These questions are naturally occurring ---they are generated in unprompted and unconstrained settings. Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context. The text-pair classification setup is similar to existing natural language inference tasks. This version uses contrast sets. These evaluation sets are expert-generated perturbations that deviate from the patterns common in the original dataset. Download size: 438.51 KiB Dataset size: 462.35 KiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 340 'validation' 340 Examples (tfds.as_dataframe): Citation: @inproceedings{clark-etal-2019-boolq, title = "{B}ool{Q}: Exploring the Surprising Difficulty of Natural Yes/No Questions", author = "Clark, Christopher and Lee, Kenton and Chang, Ming-Wei and Kwiatkowski, Tom and Collins, Michael and Toutanova, Kristina", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)", month = jun, year = "2019", address = "Minneapolis, Minnesota", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N19-1300", doi = "10.18653/v1/N19-1300", pages = "2924--2936", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/contrast_sets_drop Config description: DROP is a crowdsourced, adversarially-created QA benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs than what was necessary for prior datasets. This version uses contrast sets. These evaluation sets are expert-generated perturbations that deviate from the patterns common in the original dataset. Download size: 2.20 MiB Dataset size: 2.26 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 947 'validation' 947 Examples (tfds.as_dataframe): Citation: @inproceedings{dua-etal-2019-drop, title = "{DROP}: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs", author = "Dua, Dheeru and Wang, Yizhong and Dasigi, Pradeep and Stanovsky, Gabriel and Singh, Sameer and Gardner, Matt", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)", month = jun, year = "2019", address = "Minneapolis, Minnesota", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N19-1246", doi = "10.18653/v1/N19-1246", pages = "2368--2378", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/contrast_sets_quoref Config description: This dataset tests the coreferential reasoning capability of reading comprehension systems. In this span-selection benchmark containing questions over paragraphs from Wikipedia, a system must resolve hard coreferences before selecting the appropriate span(s) in the paragraphs for answering questions. This version uses contrast sets. These evaluation sets are expert-generated perturbations that deviate from the patterns common in the original dataset. Download size: 2.60 MiB Dataset size: 2.65 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 700 'validation' 700 Examples (tfds.as_dataframe): Citation: @inproceedings{dasigi-etal-2019-quoref, title = "{Q}uoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning", author = "Dasigi, Pradeep and Liu, Nelson F. and Marasovi{'c}, Ana and Smith, Noah A. and Gardner, Matt", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)", month = nov, year = "2019", address = "Hong Kong, China", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D19-1606", doi = "10.18653/v1/D19-1606", pages = "5925--5932", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/contrast_sets_ropes Config description: This dataset tests a system's ability to apply knowledge from a passage of text to a new situation. A system is presented a background passage containing a causal or qualitative relation(s) (e.g., "animal pollinators increase efficiency of fertilization in flowers"), a novel situation that uses this background, and questions that require reasoning about effects of the relationships in the background passage in the context of the situation. This version uses contrast sets. These evaluation sets are expert-generated perturbations that deviate from the patterns common in the original dataset. Download size: 1.97 MiB Dataset size: 2.04 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 974 'validation' 974 Examples (tfds.as_dataframe): Citation: @inproceedings{lin-etal-2019-reasoning, title = "Reasoning Over Paragraph Effects in Situations", author = "Lin, Kevin and Tafjord, Oyvind and Clark, Peter and Gardner, Matt", booktitle = "Proceedings of the 2nd Workshop on Machine Reading for Question Answering", month = nov, year = "2019", address = "Hong Kong, China", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D19-5808", doi = "10.18653/v1/D19-5808", pages = "58--62", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/drop Config description: DROP is a crowdsourced, adversarially-created QA benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs than what was necessary for prior datasets. Download size: 105.18 MiB Dataset size: 108.16 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 77,399 'validation' 9,536 Examples (tfds.as_dataframe): Citation: @inproceedings{dua-etal-2019-drop, title = "{DROP}: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs", author = "Dua, Dheeru and Wang, Yizhong and Dasigi, Pradeep and Stanovsky, Gabriel and Singh, Sameer and Gardner, Matt", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)", month = jun, year = "2019", address = "Minneapolis, Minnesota", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N19-1246", doi = "10.18653/v1/N19-1246", pages = "2368--2378", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/mctest Config description: MCTest requires machines to answer multiple-choice reading comprehension questions about fictional stories, directly tackling the high-level goal of open-domain machine comprehension. Reading comprehension can test advanced abilities such as causal reasoning and understanding the world, yet, by being multiple-choice, still provide a clear metric. By being fictional, the answer typically can be found only in the story itself. The stories and questions are also carefully limited to those a young child would understand, reducing the world knowledge that is required for the task. Download size: 2.14 MiB Dataset size: 2.20 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 1,480 'validation' 320 Examples (tfds.as_dataframe): Citation: @inproceedings{richardson-etal-2013-mctest, title = "{MCT}est: A Challenge Dataset for the Open-Domain Machine Comprehension of Text", author = "Richardson, Matthew and Burges, Christopher J.C. and Renshaw, Erin", booktitle = "Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing", month = oct, year = "2013", address = "Seattle, Washington, USA", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D13-1020", pages = "193--203", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/mctest_corrected_the_separator Config description: MCTest requires machines to answer multiple-choice reading comprehension questions about fictional stories, directly tackling the high-level goal of open-domain machine comprehension. Reading comprehension can test advanced abilities such as causal reasoning and understanding the world, yet, by being multiple-choice, still provide a clear metric. By being fictional, the answer typically can be found only in the story itself. The stories and questions are also carefully limited to those a young child would understand, reducing the world knowledge that is required for the task. Download size: 2.15 MiB Dataset size: 2.21 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 1,480 'validation' 320 Examples (tfds.as_dataframe): Citation: @inproceedings{richardson-etal-2013-mctest, title = "{MCT}est: A Challenge Dataset for the Open-Domain Machine Comprehension of Text", author = "Richardson, Matthew and Burges, Christopher J.C. and Renshaw, Erin", booktitle = "Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing", month = oct, year = "2013", address = "Seattle, Washington, USA", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D13-1020", pages = "193--203", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/multirc Config description: MultiRC is a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences. Questions and answers for this challenge were solicited and verified through a 4-step crowdsourcing experiment. The dataset contains questions for paragraphs across 7 different domains ( elementary school science, news, travel guides, fiction stories, etc) bringing in linguistic diversity to the texts and to the questions wordings. Download size: 897.09 KiB Dataset size: 918.42 KiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 312 'validation' 312 Examples (tfds.as_dataframe): Citation: @inproceedings{khashabi-etal-2018-looking, title = "Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences", author = "Khashabi, Daniel and Chaturvedi, Snigdha and Roth, Michael and Upadhyay, Shyam and Roth, Dan", booktitle = "Proceedings of the 2018 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)", month = jun, year = "2018", address = "New Orleans, Louisiana", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N18-1023", doi = "10.18653/v1/N18-1023", pages = "252--262", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/narrativeqa Config description: NarrativeQA is an English-lanaguage dataset of stories and corresponding questions designed to test reading comprehension, especially on long documents. Download size: 308.28 MiB Dataset size: 311.22 MiB Auto-cached (documentation): No Splits: Split Examples 'test' 21,114 'train' 65,494 'validation' 6,922 Examples (tfds.as_dataframe): Citation: @article{kocisky-etal-2018-narrativeqa, title = "The {N}arrative{QA} Reading Comprehension Challenge", author = "Ko{ {c} }isk{'y}, Tom{'a}{ {s} } and Schwarz, Jonathan and Blunsom, Phil and Dyer, Chris and Hermann, Karl Moritz and Melis, G{'a}bor and Grefenstette, Edward", journal = "Transactions of the Association for Computational Linguistics", volume = "6", year = "2018", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q18-1023", doi = "10.1162/tacl_a_00023", pages = "317--328", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/narrativeqa_dev Config description: NarrativeQA is an English-lanaguage dataset of stories and corresponding questions designed to test reading comprehension, especially on long documents. Download size: 308.28 MiB Dataset size: 311.22 MiB Auto-cached (documentation): No Splits: Split Examples 'test' 21,114 'train' 65,494 'validation' 6,922 Examples (tfds.as_dataframe): Citation: @article{kocisky-etal-2018-narrativeqa, title = "The {N}arrative{QA} Reading Comprehension Challenge", author = "Ko{ {c} }isk{'y}, Tom{'a}{ {s} } and Schwarz, Jonathan and Blunsom, Phil and Dyer, Chris and Hermann, Karl Moritz and Melis, G{'a}bor and Grefenstette, Edward", journal = "Transactions of the Association for Computational Linguistics", volume = "6", year = "2018", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q18-1023", doi = "10.1162/tacl_a_00023", pages = "317--328", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/natural_questions Config description: The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Download size: 6.95 MiB Dataset size: 9.88 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 96,075 'validation' 2,295 Examples (tfds.as_dataframe): Citation: @article{kwiatkowski-etal-2019-natural, title = "Natural Questions: A Benchmark for Question Answering Research", author = "Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav", journal = "Transactions of the Association for Computational Linguistics", volume = "7", year = "2019", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q19-1026", doi = "10.1162/tacl_a_00276", pages = "452--466", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/natural_questions_direct_ans Config description: The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. This version consists of direct-answer questions. Download size: 6.82 MiB Dataset size: 10.19 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 6,468 'train' 96,676 'validation' 10,693 Examples (tfds.as_dataframe): Citation: @article{kwiatkowski-etal-2019-natural, title = "Natural Questions: A Benchmark for Question Answering Research", author = "Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav", journal = "Transactions of the Association for Computational Linguistics", volume = "7", year = "2019", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q19-1026", doi = "10.1162/tacl_a_00276", pages = "452--466", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/natural_questions_direct_ans_test Config description: The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. This version consists of direct-answer questions. Download size: 6.82 MiB Dataset size: 10.19 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 6,468 'train' 96,676 'validation' 10,693 Examples (tfds.as_dataframe): Citation: @article{kwiatkowski-etal-2019-natural, title = "Natural Questions: A Benchmark for Question Answering Research", author = "Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav", journal = "Transactions of the Association for Computational Linguistics", volume = "7", year = "2019", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q19-1026", doi = "10.1162/tacl_a_00276", pages = "452--466", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/natural_questions_with_dpr_para Config description: The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. This version includes additional paragraphs (obtained using the DPR retrieval engine) to augment each question. Download size: 319.22 MiB Dataset size: 322.91 MiB Auto-cached (documentation): No Splits: Split Examples 'train' 96,676 'validation' 10,693 Examples (tfds.as_dataframe): Citation: @article{kwiatkowski-etal-2019-natural, title = "Natural Questions: A Benchmark for Question Answering Research", author = "Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav", journal = "Transactions of the Association for Computational Linguistics", volume = "7", year = "2019", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q19-1026", doi = "10.1162/tacl_a_00276", pages = "452--466", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/natural_questions_with_dpr_para_test Config description: The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. This version includes additional paragraphs (obtained using the DPR retrieval engine) to augment each question. Download size: 306.94 MiB Dataset size: 310.48 MiB Auto-cached (documentation): No Splits: Split Examples 'test' 6,468 'train' 96,676 Examples (tfds.as_dataframe): Citation: @article{kwiatkowski-etal-2019-natural, title = "Natural Questions: A Benchmark for Question Answering Research", author = "Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav", journal = "Transactions of the Association for Computational Linguistics", volume = "7", year = "2019", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/Q19-1026", doi = "10.1162/tacl_a_00276", pages = "452--466", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/newsqa Config description: NewsQA is a challenging machine comprehension dataset of human-generated question-answer pairs. Crowdworkers supply questions and answers based on a set of news articles from CNN, with answers consisting of spans of text from the corresponding articles. Download size: 283.33 MiB Dataset size: 285.94 MiB Auto-cached (documentation): No Splits: Split Examples 'train' 75,882 'validation' 4,309 Examples (tfds.as_dataframe): Citation: @inproceedings{trischler-etal-2017-newsqa, title = "{N}ews{QA}: A Machine Comprehension Dataset", author = "Trischler, Adam and Wang, Tong and Yuan, Xingdi and Harris, Justin and Sordoni, Alessandro and Bachman, Philip and Suleman, Kaheer", booktitle = "Proceedings of the 2nd Workshop on Representation Learning for {NLP}", month = aug, year = "2017", address = "Vancouver, Canada", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/W17-2623", doi = "10.18653/v1/W17-2623", pages = "191--200", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/openbookqa Config description: OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject. Download size: 942.34 KiB Dataset size: 1.11 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 500 'train' 4,957 'validation' 500 Examples (tfds.as_dataframe): Citation: @inproceedings{mihaylov-etal-2018-suit, title = "Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering", author = "Mihaylov, Todor and Clark, Peter and Khot, Tushar and Sabharwal, Ashish", booktitle = "Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing", month = oct # "-" # nov, year = "2018", address = "Brussels, Belgium", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D18-1260", doi = "10.18653/v1/D18-1260", pages = "2381--2391", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/openbookqa_dev Config description: OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject. Download size: 942.34 KiB Dataset size: 1.11 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 500 'train' 4,957 'validation' 500 Examples (tfds.as_dataframe): Citation: @inproceedings{mihaylov-etal-2018-suit, title = "Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering", author = "Mihaylov, Todor and Clark, Peter and Khot, Tushar and Sabharwal, Ashish", booktitle = "Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing", month = oct # "-" # nov, year = "2018", address = "Brussels, Belgium", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D18-1260", doi = "10.18653/v1/D18-1260", pages = "2381--2391", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/openbookqa_with_ir Config description: OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 6.08 MiB Dataset size: 6.28 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 500 'train' 4,957 'validation' 500 Examples (tfds.as_dataframe): Citation: @inproceedings{mihaylov-etal-2018-suit, title = "Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering", author = "Mihaylov, Todor and Clark, Peter and Khot, Tushar and Sabharwal, Ashish", booktitle = "Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing", month = oct # "-" # nov, year = "2018", address = "Brussels, Belgium", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D18-1260", doi = "10.18653/v1/D18-1260", pages = "2381--2391", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/openbookqa_with_ir_dev Config description: OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 6.08 MiB Dataset size: 6.28 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 500 'train' 4,957 'validation' 500 Examples (tfds.as_dataframe): Citation: @inproceedings{mihaylov-etal-2018-suit, title = "Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering", author = "Mihaylov, Todor and Clark, Peter and Khot, Tushar and Sabharwal, Ashish", booktitle = "Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing", month = oct # "-" # nov, year = "2018", address = "Brussels, Belgium", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D18-1260", doi = "10.18653/v1/D18-1260", pages = "2381--2391", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/physical_iqa Config description: This is a dataset for benchmarking progress in physical commonsense understanding. The underlying task is multiple choice question answering: given a question q and two possible solutions s1, s2, a model or a human must choose the most appropriate solution, of which exactly one is correct. The dataset focuses on everyday situations with a preference for atypical solutions. The dataset is inspired by instructables.com, which provides users with instructions on how to build, craft, bake, or manipulate objects using everyday materials. Annotators are asked to provide semantic perturbations or alternative approaches which are otherwise syntactically and topically similar to ensure physical knowledge is targeted. The dataset is further cleaned of basic artifacts using the AFLite algorithm. Download size: 6.01 MiB Dataset size: 6.59 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 16,113 'validation' 1,838 Examples (tfds.as_dataframe): Citation: @inproceedings{bisk2020piqa, title={Piqa: Reasoning about physical commonsense in natural language}, author={Bisk, Yonatan and Zellers, Rowan and Gao, Jianfeng and Choi, Yejin and others}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={7432--7439}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/qasc Config description: QASC is a question-answering dataset with a focus on sentence composition. It consists of 8-way multiple-choice questions about grade school science, and comes with a corpus of 17M sentences. Download size: 1.75 MiB Dataset size: 2.09 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 920 'train' 8,134 'validation' 926 Examples (tfds.as_dataframe): Citation: @inproceedings{khot2020qasc, title={Qasc: A dataset for question answering via sentence composition}, author={Khot, Tushar and Clark, Peter and Guerquin, Michal and Jansen, Peter and Sabharwal, Ashish}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8082--8090}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/qasc_test Config description: QASC is a question-answering dataset with a focus on sentence composition. It consists of 8-way multiple-choice questions about grade school science, and comes with a corpus of 17M sentences. Download size: 1.75 MiB Dataset size: 2.09 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 920 'train' 8,134 'validation' 926 Examples (tfds.as_dataframe): Citation: @inproceedings{khot2020qasc, title={Qasc: A dataset for question answering via sentence composition}, author={Khot, Tushar and Clark, Peter and Guerquin, Michal and Jansen, Peter and Sabharwal, Ashish}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8082--8090}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/qasc_with_ir Config description: QASC is a question-answering dataset with a focus on sentence composition. It consists of 8-way multiple-choice questions about grade school science, and comes with a corpus of 17M sentences. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 16.95 MiB Dataset size: 17.30 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 920 'train' 8,134 'validation' 926 Examples (tfds.as_dataframe): Citation: @inproceedings{khot2020qasc, title={Qasc: A dataset for question answering via sentence composition}, author={Khot, Tushar and Clark, Peter and Guerquin, Michal and Jansen, Peter and Sabharwal, Ashish}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8082--8090}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/qasc_with_ir_test Config description: QASC is a question-answering dataset with a focus on sentence composition. It consists of 8-way multiple-choice questions about grade school science, and comes with a corpus of 17M sentences. This version includes paragraphs fetched via an information retrieval system as additional evidence. Download size: 16.95 MiB Dataset size: 17.30 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 920 'train' 8,134 'validation' 926 Examples (tfds.as_dataframe): Citation: @inproceedings{khot2020qasc, title={Qasc: A dataset for question answering via sentence composition}, author={Khot, Tushar and Clark, Peter and Guerquin, Michal and Jansen, Peter and Sabharwal, Ashish}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8082--8090}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/quoref Config description: This dataset tests the coreferential reasoning capability of reading comprehension systems. In this span-selection benchmark containing questions over paragraphs from Wikipedia, a system must resolve hard coreferences before selecting the appropriate span(s) in the paragraphs for answering questions. Download size: 51.43 MiB Dataset size: 52.29 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 22,265 'validation' 2,768 Examples (tfds.as_dataframe): Citation: @inproceedings{dasigi-etal-2019-quoref, title = "{Q}uoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning", author = "Dasigi, Pradeep and Liu, Nelson F. and Marasovi{'c}, Ana and Smith, Noah A. and Gardner, Matt", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)", month = nov, year = "2019", address = "Hong Kong, China", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D19-1606", doi = "10.18653/v1/D19-1606", pages = "5925--5932", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/race_string Config description: Race is a large-scale reading comprehension dataset. The dataset is collected from English examinations in China, which are designed for middle school and high school students. The dataset can be served as the training and test sets for machine comprehension. Download size: 167.97 MiB Dataset size: 171.23 MiB Auto-cached (documentation): Yes (test, validation), Only when shuffle_files=False (train) Splits: Split Examples 'test' 4,934 'train' 87,863 'validation' 4,887 Examples (tfds.as_dataframe): Citation: @inproceedings{lai-etal-2017-race, title = "{RACE}: Large-scale {R}e{A}ding Comprehension Dataset From Examinations", author = "Lai, Guokun and Xie, Qizhe and Liu, Hanxiao and Yang, Yiming and Hovy, Eduard", booktitle = "Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing", month = sep, year = "2017", address = "Copenhagen, Denmark", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D17-1082", doi = "10.18653/v1/D17-1082", pages = "785--794", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/race_string_dev Config description: Race is a large-scale reading comprehension dataset. The dataset is collected from English examinations in China, which are designed for middle school and high school students. The dataset can be served as the training and test sets for machine comprehension. Download size: 167.97 MiB Dataset size: 171.23 MiB Auto-cached (documentation): Yes (test, validation), Only when shuffle_files=False (train) Splits: Split Examples 'test' 4,934 'train' 87,863 'validation' 4,887 Examples (tfds.as_dataframe): Citation: @inproceedings{lai-etal-2017-race, title = "{RACE}: Large-scale {R}e{A}ding Comprehension Dataset From Examinations", author = "Lai, Guokun and Xie, Qizhe and Liu, Hanxiao and Yang, Yiming and Hovy, Eduard", booktitle = "Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing", month = sep, year = "2017", address = "Copenhagen, Denmark", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D17-1082", doi = "10.18653/v1/D17-1082", pages = "785--794", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/ropes Config description: This dataset tests a system's ability to apply knowledge from a passage of text to a new situation. A system is presented a background passage containing a causal or qualitative relation(s) (e.g., "animal pollinators increase efficiency of fertilization in flowers"), a novel situation that uses this background, and questions that require reasoning about effects of the relationships in the background passage in the context of the situation. Download size: 12.91 MiB Dataset size: 13.35 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 10,924 'validation' 1,688 Examples (tfds.as_dataframe): Citation: @inproceedings{lin-etal-2019-reasoning, title = "Reasoning Over Paragraph Effects in Situations", author = "Lin, Kevin and Tafjord, Oyvind and Clark, Peter and Gardner, Matt", booktitle = "Proceedings of the 2nd Workshop on Machine Reading for Question Answering", month = nov, year = "2019", address = "Hong Kong, China", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D19-5808", doi = "10.18653/v1/D19-5808", pages = "58--62", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/social_iqa Config description: This is a large-scale benchmark for commonsense reasoning about social situations. Social IQa contains multiple choice questions for probing emotional and social intelligence in a variety of everyday situations. Through crowdsourcing, commonsense questions along with correct and incorrect answers about social interactions are collected, using a new framework that mitigates stylistic artifacts in incorrect answers by asking workers to provide the right answer to a different but related question. Download size: 7.08 MiB Dataset size: 8.22 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 33,410 'validation' 1,954 Examples (tfds.as_dataframe): Citation: @inproceedings{sap-etal-2019-social, title = "Social {IQ}a: Commonsense Reasoning about Social Interactions", author = "Sap, Maarten and Rashkin, Hannah and Chen, Derek and Le Bras, Ronan and Choi, Yejin", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)", month = nov, year = "2019", address = "Hong Kong, China", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D19-1454", doi = "10.18653/v1/D19-1454", pages = "4463--4473", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/squad1_1 Config description: This is a reading comprehension dataset consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage. Download size: 80.62 MiB Dataset size: 83.99 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 87,514 'validation' 10,570 Examples (tfds.as_dataframe): Citation: @inproceedings{rajpurkar-etal-2016-squad, title = "{SQ}u{AD}: 100,000+ Questions for Machine Comprehension of Text", author = "Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy", booktitle = "Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing", month = nov, year = "2016", address = "Austin, Texas", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/D16-1264", doi = "10.18653/v1/D16-1264", pages = "2383--2392", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/squad2 Config description: This dataset combines the original Stanford Question Answering Dataset (SQuAD) dataset with unanswerable questions written adversarially by crowdworkers to look similar to answerable ones. Download size: 116.56 MiB Dataset size: 121.43 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 130,149 'validation' 11,873 Examples (tfds.as_dataframe): Citation: @inproceedings{rajpurkar-etal-2018-know, title = "Know What You Don{'}t Know: Unanswerable Questions for {SQ}u{AD}", author = "Rajpurkar, Pranav and Jia, Robin and Liang, Percy", booktitle = "Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)", month = jul, year = "2018", address = "Melbourne, Australia", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/P18-2124", doi = "10.18653/v1/P18-2124", pages = "784--789", } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/winogrande_l Config description: This dataset is inspired by the original Winograd Schema Challenge design, but adjusted to improve both the scale and the hardness of the dataset. The key steps of the dataset construction consist of (1) a carefully designed crowdsourcing procedure, followed by (2) systematic bias reduction using a novel AfLite algorithm that generalizes human-detectable word associations to machine-detectable embedding associations. Training sets with differnt sizes are provided. This set corresponds to size l. Download size: 1.49 MiB Dataset size: 1.83 MiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 10,234 'validation' 1,267 Examples (tfds.as_dataframe): Citation: @inproceedings{sakaguchi2020winogrande, title={Winogrande: An adversarial winograd schema challenge at scale}, author={Sakaguchi, Keisuke and Le Bras, Ronan and Bhagavatula, Chandra and Choi, Yejin}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8732--8740}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/winogrande_m Config description: This dataset is inspired by the original Winograd Schema Challenge design, but adjusted to improve both the scale and the hardness of the dataset. The key steps of the dataset construction consist of (1) a carefully designed crowdsourcing procedure, followed by (2) systematic bias reduction using a novel AfLite algorithm that generalizes human-detectable word associations to machine-detectable embedding associations. Training sets with differnt sizes are provided. This set corresponds to size m. Download size: 507.46 KiB Dataset size: 623.15 KiB Auto-cached (documentation): Yes Splits: Split Examples 'train' 2,558 'validation' 1,267 Examples (tfds.as_dataframe): Citation: @inproceedings{sakaguchi2020winogrande, title={Winogrande: An adversarial winograd schema challenge at scale}, author={Sakaguchi, Keisuke and Le Bras, Ronan and Bhagavatula, Chandra and Choi, Yejin}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8732--8740}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." unified_qa/winogrande_s Config description: This dataset is inspired by the original Winograd Schema Challenge design, but adjusted to improve both the scale and the hardness of the dataset. The key steps of the dataset construction consist of (1) a carefully designed crowdsourcing procedure, followed by (2) systematic bias reduction using a novel AfLite algorithm that generalizes human-detectable word associations to machine-detectable embedding associations. Training sets with differnt sizes are provided. This set corresponds to size s. Download size: 479.24 KiB Dataset size: 590.47 KiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 1,767 'train' 640 'validation' 1,267 Examples (tfds.as_dataframe): Citation: @inproceedings{sakaguchi2020winogrande, title={Winogrande: An adversarial winograd schema challenge at scale}, author={Sakaguchi, Keisuke and Le Bras, Ronan and Bhagavatula, Chandra and Choi, Yejin}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={34}, number={05}, pages={8732--8740}, year={2020} } @inproceedings{khashabi-etal-2020-unifiedqa, title = "{UNIFIEDQA}: Crossing Format Boundaries with a Single {QA} System", author = "Khashabi, Daniel and Min, Sewon and Khot, Tushar and Sabharwal, Ashish and Tafjord, Oyvind and Clark, Peter and Hajishirzi, Hannaneh", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2020", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.findings-emnlp.171", doi = "10.18653/v1/2020.findings-emnlp.171", pages = "1896--1907", } Note that each UnifiedQA dataset has its own citation. Please see the source to see the correct citation for each contained dataset." Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License, and code samples are licensed under the Apache 2.0 License. For details, see the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates. Last updated 2022-12-06 UTC. [[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Missing the information I need","missingTheInformationINeed","thumb-down"],["Too complicated / too many steps","tooComplicatedTooManySteps","thumb-down"],["Out of date","outOfDate","thumb-down"],["Samples / code issue","samplesCodeIssue","thumb-down"],["Other","otherDown","thumb-down"]],["Last updated 2022-12-06 UTC."],[],[]] Stay connected Blog Forum GitHub Twitter YouTube Support Issue tracker Release notes Stack Overflow Brand guidelines Cite TensorFlow Terms Privacy Manage cookies Sign up for the TensorFlow newsletter Subscribe English Español – América Latina Français Indonesia Italiano Polski Português – Brasil Tiếng Việt Türkçe Русский עברית العربيّة فارسی हिंदी বাংলা ภาษาไทย 中文 – 简体 日本語 한국어