wmt18_translate  |  TensorFlow Datasets Skip to main content Install Learn Introduction New to TensorFlow? Tutorials Learn how to use TensorFlow with end-to-end examples Guide Learn framework concepts and components Learn ML Educational resources to master your path with TensorFlow API TensorFlow (v2.16.1) Versions… TensorFlow.js TensorFlow Lite TFX Resources LIBRARIES TensorFlow.js Develop web ML applications in JavaScript TensorFlow Lite Deploy ML on mobile, microcontrollers and other edge devices TFX Build production ML pipelines All libraries Create advanced models and extend TensorFlow RESOURCES Models & datasets Pre-trained models and datasets built by Google and the community Tools Tools to support and accelerate TensorFlow workflows Responsible AI Resources for every stage of the ML workflow Recommendation systems Build recommendation systems with open source tools Community Groups User groups, interest groups and mailing lists Contribute Guide for contributing to code and documentation Blog Stay up to date with all things TensorFlow Forum Discussion platform for the TensorFlow community Why TensorFlow About Case studies / English Español – América Latina Français Indonesia Italiano Polski Português – Brasil Tiếng Việt Türkçe Русский עברית العربيّة فارسی हिंदी বাংলা ภาษาไทย 中文 – 简体 日本語 한국어 GitHub Sign in Datasets Overview Catalog Community Catalog Guide API Install Learn More API More Resources More Overview Catalog Community Catalog Guide API Community More Why TensorFlow More GitHub Overview Dataset Collections longt5 xtreme 3d aflw2k3d smallnorb smartwatch_gestures Abstractive text summarization aeslc billsum booksum (manual) multi_news newsroom (manual) reddit reddit_tifu samsum (manual) scientific_papers Age dices wake_vision Anomaly detection ag_news_subset caltech101 kddcup99 lost_and_found stl10 Audio accentdb common_voice crema_d dementiabank (manual) fuss groove gtzan gtzan_music_speech librispeech libritts ljspeech nsynth savee (manual) speech_commands spoken_digit tedlium user_libri_audio vctk voxceleb (manual) voxforge (manual) xtreme_s yes_no Biology ai2dcaption ogbg_molpcba Categorical dices sift1m wake_vision Common sense reasoning ai2_arc_with_ir arc covr natural_questions openbookqa Computer science robomimic_mg robomimic_mh robomimic_ph smart_buildings Conditional image generation imagenet2012 (manual) imagenet2012_subset (manual) webvid (manual) Coreference resolution clevr D4rl d4rl_adroit_door d4rl_adroit_hammer d4rl_adroit_pen d4rl_adroit_relocate d4rl_antmaze d4rl_mujoco_ant d4rl_mujoco_halfcheetah d4rl_mujoco_hopper d4rl_mujoco_walker2d Density estimation caltech101 celeb_a_hq (manual) imagenet2012 (manual) imagenet2012_subset (manual) Dependency parsing universal_dependencies xtreme_pos Dialog act labeling bot_adversarial_dialogue Dialogue bot_adversarial_dialogue databricks_dolly dices Document summarization aeslc booksum (manual) newsroom (manual) scientific_papers Facial attributes wake_vision Fine grained image classification caltech101 oxford_flowers102 oxford_iiit_pet stanford_dogs stl10 sun397 wake_vision Gender dices wake_vision Graph ogbg_molpcba reddit Graphs cardiotox Health pneumonia_mnist Image abstract_reasoning (manual) aflw2k3d ai2dcaption bccd beans bee_dataset bigearthnet binarized_mnist binary_alpha_digits caltech101 celeb_a celeb_a_hq (manual) cityscapes (manual) clevr clic coil100 covr div2k downsampled_imagenet dsprites flic imagenet2012 (manual) imagenet2012_corrupted (manual) imagenet2012_fewshot (manual) imagenet2012_multilabel (manual) imagenet2012_real (manual) imagenet2012_subset (manual) imagenet_a imagenet_lt (manual) imagenet_pi (manual) imagenet_r imagenet_resized imagenet_sketch imagenet_v2 imagenette imagewang kitti lfw lost_and_found lsun lvis malaria nyu_depth_v2 open_images_challenge2019_detection open_images_v4 oxford_flowers102 oxford_iiit_pet pass patch_camelyon pet_finder places365_small placesfull plant_leaves plant_village plantae_k pneumonia_mnist quickdraw_bitmap ref_coco (manual) resisc45 (manual) robomimic_mg robomimic_mh robomimic_ph rock_paper_scissors s3o4d scene_parse150 shapes3d siscore smallnorb so2sat stanford_dogs stanford_online_products stl10 sun397 svhn_cropped symmetric_solids tf_flowers the300w_lp wake_vision Image classification abstract_reasoning (manual) bigearthnet caltech101 caltech_birds2010 caltech_birds2011 cars196 cassava cats_vs_dogs celeb_a chexpert (manual) cifar10 cifar100 cifar100_n (manual) cifar10_1 cifar10_corrupted cifar10_h cifar10_n (manual) citrus_leaves cmaterdb colorectal_histology colorectal_histology_large controlled_noisy_web_labels (manual) curated_breast_imaging_ddsm (manual) cycle_gan deep_weeds diabetic_retinopathy_detection (manual) dmlab domainnet dtd emnist eurosat fashion_mnist flic food101 geirhos_conflict_stimuli horses_or_humans i_naturalist2017 i_naturalist2018 i_naturalist2021 imagenet2012 (manual) imagenet2012_subset (manual) imagenet_resized imagenet_sketch imagenette imagewang kitti kmnist lvis malaria mnist mnist_corrupted omniglot open_images_v4 oxford_flowers102 oxford_iiit_pet places365_small plant_village pneumonia_mnist resisc45 (manual) siscore smallnorb stanford_dogs stanford_online_products stl10 sun397 svhn_cropped uc_merced visual_domain_decathlon wake_vision Image clustering imagenet2012 (manual) imagenet2012_subset (manual) stanford_dogs stl10 Image compression imagenet2012 (manual) imagenet2012_subset (manual) imagenet_resized oxford_iiit_pet patch_camelyon stl10 Image generation binarized_mnist celeb_a celeb_a_hq (manual) cityscapes (manual) clevr imagenet2012 (manual) imagenet2012_subset (manual) oxford_flowers102 stanford_dogs stl10 Image segmentation segment_anything (manual) Image super resolution celeb_a_hq (manual) div2k kitti Image to image translation celeb_a_hq (manual) cityscapes (manual) kitti scene_parse150 Instance segmentation cityscapes (manual) lvis nyu_depth_v2 segment_anything (manual) Language modeling big_patent billsum blimp databricks_dolly dices dolma e2e_cleaned irc_disentanglement lambada librispeech_lm lm1b math_qa opus paws_wiki paws_x_wiki pg19 reddit samsum (manual) snli squad tedlium Linguistic acceptability bot_adversarial_dialogue Machine translation mlqa opus Monolingual ag_news_subset ai2_arc_with_ir arc beir booksum (manual) bool_q bot_adversarial_dialogue covr dices dolma e2e_cleaned imdb_reviews kitti lambada librispeech librispeech_lm libritts ljspeech lm1b natural_questions natural_questions_open openbookqa paws_wiki plant_village quac race real_toxicity_prompts reddit savee (manual) schema_guided_dialogue sci_tail scicite scientific_papers sentiment140 snli speech_commands spoken_digit squad story_cloze (manual) tedlium trec trivia_qa Movies and tv shows opinion_abstracts Multilingual librispeech mlqa paws_x_wiki Natural language inference anli covr dices paws_wiki sci_tail snli Natural language understanding ag_news_subset anli beir bool_q clevr covr databricks_dolly dices imdb_reviews math_dataset math_qa mlqa multi_news natural_instructions natural_questions natural_questions_open openbookqa opus paws_wiki paws_x_wiki pg19 piqa qasc quac race schema_guided_dialogue sci_tail sentiment140 snli squad story_cloze (manual) trec trivia_qa Nearest neighbors deep1b glove100_angular News multi_news Object detection coco coco_captions covr flic kitti lvis open_images_v4 voc waymo_open_dataset wider_face Open domain question answering databricks_dolly natural_questions squad trivia_qa Out of distribution detection stl10 Question answering beir bool_q clevr coqa cosmos_qa databricks_dolly math_dataset math_qa mlqa natural_instructions natural_questions natural_questions_open openbookqa piqa qasc quac race squad story_cloze (manual) trivia_qa tydi_qa web_questions xquad Question generation natural_questions trivia_qa Race national ethnic origin dices Ranking istella mslr_web yahoo_ltrc (manual) Reading comprehension opus pg19 qasc quac race squad trivia_qa Recommendation criteo hillstrom Reinforcement learning robomimic_mg robomimic_mh robomimic_ph smart_buildings Rgb d nyu_depth_v2 Rl unplugged rlu_atari rlu_atari_checkpoints rlu_atari_checkpoints_ordered rlu_control_suite rlu_dmlab_explore_object_rewards_few rlu_dmlab_explore_object_rewards_many rlu_dmlab_rooms_select_nonmatching_object rlu_dmlab_rooms_watermaze rlu_dmlab_seekavoid_arena01 rlu_locomotion rlu_rwrl Rlds locomotion robosuite_panda_pick_place_can Robotics aloha_mobile asimov_dilemmas_auto_val asimov_dilemmas_scifi_train asimov_dilemmas_scifi_val asimov_injury_val asimov_multimodal_auto_val asimov_multimodal_manual_val asimov_v2_constraints_with_rationale asimov_v2_constraints_without_rationale asimov_v2_injuries asimov_v2_videos asu_table_top_converted_externally_to_rlds austin_buds_dataset_converted_externally_to_rlds austin_sailor_dataset_converted_externally_to_rlds austin_sirius_dataset_converted_externally_to_rlds bc_z berkeley_autolab_ur5 berkeley_cable_routing berkeley_fanuc_manipulation berkeley_gnm_cory_hall berkeley_gnm_recon berkeley_gnm_sac_son berkeley_mvp_converted_externally_to_rlds berkeley_rpt_converted_externally_to_rlds bridge bridge_data_msr cmu_franka_exploration_dataset_converted_externally_to_rlds cmu_play_fusion cmu_stretch columbia_cairlab_pusht_real conq_hose_manipulation dlr_edan_shared_control_converted_externally_to_rlds dlr_sara_grid_clamp_converted_externally_to_rlds dlr_sara_pour_converted_externally_to_rlds dobbe eth_agent_affordances fmb fractal20220817_data iamlab_cmu_pickup_insert_converted_externally_to_rlds imperialcollege_sawyer_wrist_cam io_ai_tech jaco_play kaist_nonprehensile_converted_externally_to_rlds kuka maniskill_dataset_converted_externally_to_rlds mimic_play mt_opt nyu_door_opening_surprising_effectiveness nyu_franka_play_dataset_converted_externally_to_rlds nyu_rot_dataset_converted_externally_to_rlds plex_robosuite robo_ai_u_r5e robo_set roboturk spoc_robot stanford_hydra_dataset_converted_externally_to_rlds stanford_kuka_multimodal_dataset_converted_externally_to_rlds stanford_mask_vit_converted_externally_to_rlds stanford_robocook_converted_externally_to_rlds taco_play tidybot tokyo_u_lsmo_converted_externally_to_rlds toto ucsd_kitchen_dataset_converted_externally_to_rlds ucsd_pick_and_place_dataset_converted_externally_to_rlds uiuc_d3field usc_cloth_sim_converted_externally_to_rlds utaustin_mutex utokyo_pr2_opening_fridge_converted_externally_to_rlds utokyo_pr2_tabletop_manipulation_converted_externally_to_rlds utokyo_saytap_converted_externally_to_rlds utokyo_xarm_bimanual_converted_externally_to_rlds utokyo_xarm_pick_and_place_converted_externally_to_rlds vima_converted_externally_to_rlds viola Scene classification bigearthnet places365_small resisc45 (manual) Semantic segmentation bigearthnet cityscapes (manual) kitti lost_and_found nyu_depth_v2 open_images_v4 places365_small ref_coco (manual) scene_parse150 segment_anything (manual) so2sat Sentiment analysis imdb_reviews sentiment140 Sequence modeling databricks_dolly smart_buildings Sequence to sequence language modeling big_patent billsum databricks_dolly math_qa opus paws_wiki reddit samsum (manual) snli Speech librispeech libritts speech_commands Speech recognition accentdb librispeech speech_commands tedlium Structured cherry_blossoms covid19 cs_restaurants dart diamonds forest_fires genomics_ood german_credit_numeric higgs howell iris movie_lens movielens web_graph web_nlg wiki_bio wiki_table_questions wiki_table_text wine_quality Summarization cnn_dailymail covid19sum (manual) gigaword gov_report wikihow (manual) xsum (manual) Table to text generation e2e_cleaned Tabular ble_wind_field efron_morris75 kddcup99 opinion_abstracts radon simpte (manual) titanic Text abstract_reasoning (manual) aeslc ag_news_subset ai2_arc ai2_arc_with_ir amazon_us_reviews anli answer_equivalence arc asqa asset assin2 bccd beir big_patent billsum blimp booksum (manual) bool_q bot_adversarial_dialogue bucc c4 (manual) c4_wsrs caltech101 cfq cityscapes (manual) civil_comments clevr clinc_oos conll2002 conll2003 corr2cause cos_e covr databricks_dolly definite_pronoun_resolution dices doc_nli dolma dolphin_number_word drop dsprites e2e_cleaned eraser_multi_rc esnli flic gap gem glue goemotions gpt3 gsm8k hellaswag imdb_reviews irc_disentanglement kitti lambada lfw librispeech librispeech_lm libritts ljspeech lm1b lost_and_found math_dataset math_qa mctaco media_sum (manual) mlqa movie_rationales mrqa multi_news multi_nli multi_nli_mismatch natural_instructions natural_questions natural_questions_open newsroom (manual) open_images_challenge2019_detection open_images_v4 openbookqa opinion_abstracts opinosis opus oxford_flowers102 oxford_iiit_pet para_crawl patch_camelyon paws_wiki paws_x_wiki penguins pet_finder pg19 piqa places365_small placesfull plant_leaves plant_village plantae_k protein_net q_re_cc qa4mre qasc quac quality race real_toxicity_prompts reddit_disentanglement (manual) reddit_tifu ref_coco (manual) resisc45 (manual) robonet rock_you salient_span_wikipedia samsum (manual) scan schema_guided_dialogue sci_tail scicite scientific_papers scrolls sentiment140 smallnorb snli spoken_digit squad squad_question_generation stanford_dogs star_cfq story_cloze (manual) summscreen sun397 super_glue svhn_cropped tatoeba ted_hrlr_translate ted_multi_translate tedlium tiny_shakespeare trec trivia_qa unified_qa universal_dependencies unnatural_instructions user_libri_text webvid (manual) wiki40b wiki_dialog wikiann wikipedia wikipedia_toxicity_subtypes winogrande wordnet wsc273 xnli xtreme_pawsx xtreme_xnli yelp_polarity_reviews Text classification ag_news_subset bool_q bot_adversarial_dialogue dices imdb_reviews natural_instructions paws_wiki paws_x_wiki sentiment140 trec Text classification toxicity prediction bot_adversarial_dialogue dices real_toxicity_prompts Text generation aeslc big_patent billsum booksum (manual) bool_q databricks_dolly e2e_cleaned lambada lm1b math_qa mctaco natural_instructions natural_questions newsroom (manual) openbookqa piqa race real_toxicity_prompts reddit reddit_tifu samsum (manual) schema_guided_dialogue scientific_papers squad story_cloze (manual) trivia_qa Text simplification wiki_auto Text summarization aeslc big_patent billsum booksum (manual) databricks_dolly multi_news newsroom (manual) reddit reddit_tifu samsum (manual) scientific_papers Time series robomimic_mg robomimic_mh robomimic_ph smart_buildings smartwatch_gestures Token classification universal_dependencies xtreme_pos Tracking smartwatch_gestures Trajectory robomimic_mg robomimic_mh robomimic_ph Translate flores mtnt wmt13_translate (manual) wmt14_translate (manual) wmt15_translate (manual) wmt16_translate (manual) wmt17_translate (manual) wmt18_translate (manual) wmt19_translate (manual) wmt_t2t_translate (manual) Uncategorized duke_ultrasound lbpp qm9 Unsupervised anomaly detection caltech101 Video abstract_reasoning (manual) bair_robot_pushing_small davis flic moving_mnist robonet starcraft_video tao (manual) ucf101 webvid (manual) youtube_vis (manual) Vision language gref (manual) grounded_scan laion400m (manual) wit wit_kaggle (manual) Introduction Tutorials Guide Learn ML TensorFlow (v2.16.1) Versions… TensorFlow.js TensorFlow Lite TFX LIBRARIES TensorFlow.js TensorFlow Lite TFX All libraries RESOURCES Models & datasets Tools Responsible AI Recommendation systems Groups Contribute Blog Forum About Case studies TFDS now supports the Croissant 🥐 format! Read the documentation to know more. TensorFlow Resources Datasets Catalog wmt18_translate Stay organized with collections Save and categorize content based on your preferences. Warning: Manual download required. See instructions below. Description: Translate dataset based on the data from statmt.org. Versions exists for the different years using a combination of multiple data sources. The base wmt_translate allows you to create your own config to choose your own data/language pair by creating a custom tfds.translate.wmt.WmtConfig. config = tfds.translate.wmt.WmtConfig( version="0.0.1", language_pair=("fr", "de"), subsets={ tfds.Split.TRAIN: ["commoncrawl_frde"], tfds.Split.VALIDATION: ["euelections_dev2019"], }, ) builder = tfds.builder("wmt_translate", config=config) Additional Documentation: Explore on Papers With Code north_east Homepage: http://www.statmt.org/wmt18/translation-task.html Source code: tfds.translate.Wmt18Translate Versions: 1.0.0 (default): No release notes. Manual download instructions: This dataset requires you to download the source data manually into download_config.manual_dir (defaults to ~/tensorflow_datasets/downloads/manual/): Some of the wmt configs here, require a manual download. Please look into wmt.py to see the exact path (and file name) that has to be downloaded. Figure (tfds.show_examples): Not supported. Citation: @InProceedings{bojar-EtAl:2018:WMT1, author = {Bojar, Ond {r}ej and Federmann, Christian and Fishel, Mark and Graham, Yvette and Haddow, Barry and Huck, Matthias and Koehn, Philipp and Monz, Christof}, title = {Findings of the 2018 Conference on Machine Translation (WMT18)}, booktitle = {Proceedings of the Third Conference on Machine Translation, Volume 2: Shared Task Papers}, month = {October}, year = {2018}, address = {Belgium, Brussels}, publisher = {Association for Computational Linguistics}, pages = {272--307}, url = {http://www.aclweb.org/anthology/W18-6401} } wmt18_translate/cs-en (default config) Config description: WMT 2018 cs-en translation task dataset. Download size: 1.89 GiB Dataset size: 3.84 GiB Auto-cached (documentation): No Splits: Split Examples 'test' 2,983 'train' 24,021,877 'validation' 3,005 Feature structure: Translation({ 'cs': Text(shape=(), dtype=string), 'en': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation cs Text string en Text string Supervised keys (See as_supervised doc): ('cs', 'en') Examples (tfds.as_dataframe): wmt18_translate/de-en Config description: WMT 2018 de-en translation task dataset. Download size: 3.55 GiB Dataset size: 8.44 GiB Auto-cached (documentation): No Splits: Split Examples 'test' 2,998 'train' 42,271,874 'validation' 3,004 Feature structure: Translation({ 'de': Text(shape=(), dtype=string), 'en': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation de Text string en Text string Supervised keys (See as_supervised doc): ('de', 'en') Examples (tfds.as_dataframe): wmt18_translate/et-en Config description: WMT 2018 et-en translation task dataset. Download size: 499.91 MiB Dataset size: 663.80 MiB Auto-cached (documentation): No Splits: Split Examples 'test' 2,000 'train' 2,175,873 'validation' 2,000 Feature structure: Translation({ 'en': Text(shape=(), dtype=string), 'et': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation en Text string et Text string Supervised keys (See as_supervised doc): ('et', 'en') Examples (tfds.as_dataframe): wmt18_translate/fi-en Config description: WMT 2018 fi-en translation task dataset. Download size: 468.76 MiB Dataset size: 889.40 MiB Auto-cached (documentation): No Splits: Split Examples 'test' 3,000 'train' 3,280,600 'validation' 6,004 Feature structure: Translation({ 'en': Text(shape=(), dtype=string), 'fi': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation en Text string fi Text string Supervised keys (See as_supervised doc): ('fi', 'en') Examples (tfds.as_dataframe): wmt18_translate/ru-en Config description: WMT 2018 ru-en translation task dataset. Download size: 1.63 GiB Dataset size: 13.89 GiB Auto-cached (documentation): No Splits: Split Examples 'test' 3,000 'train' 37,858,512 'validation' 3,001 Feature structure: Translation({ 'en': Text(shape=(), dtype=string), 'ru': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation en Text string ru Text string Supervised keys (See as_supervised doc): ('ru', 'en') Examples (tfds.as_dataframe): wmt18_translate/tr-en Config description: WMT 2018 tr-en translation task dataset. Download size: 59.32 MiB Dataset size: 63.78 MiB Auto-cached (documentation): Yes Splits: Split Examples 'test' 3,000 'train' 205,756 'validation' 3,007 Feature structure: Translation({ 'en': Text(shape=(), dtype=string), 'tr': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation en Text string tr Text string Supervised keys (See as_supervised doc): ('tr', 'en') Examples (tfds.as_dataframe): wmt18_translate/zh-en Config description: WMT 2018 zh-en translation task dataset. Download size: 831.45 MiB Dataset size: 6.43 GiB Auto-cached (documentation): No Splits: Split Examples 'test' 3,981 'train' 25,162,209 'validation' 2,001 Feature structure: Translation({ 'en': Text(shape=(), dtype=string), 'zh': Text(shape=(), dtype=string), }) Feature documentation: Feature Class Shape Dtype Description Translation en Text string zh Text string Supervised keys (See as_supervised doc): ('zh', 'en') Examples (tfds.as_dataframe): Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License, and code samples are licensed under the Apache 2.0 License. For details, see the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates. Last updated 2022-12-06 UTC. [[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Missing the information I need","missingTheInformationINeed","thumb-down"],["Too complicated / too many steps","tooComplicatedTooManySteps","thumb-down"],["Out of date","outOfDate","thumb-down"],["Samples / code issue","samplesCodeIssue","thumb-down"],["Other","otherDown","thumb-down"]],["Last updated 2022-12-06 UTC."],[],[]] Stay connected Blog Forum GitHub Twitter YouTube Support Issue tracker Release notes Stack Overflow Brand guidelines Cite TensorFlow Terms Privacy Manage cookies Sign up for the TensorFlow newsletter Subscribe English Español – América Latina Français Indonesia Italiano Polski Português – Brasil Tiếng Việt Türkçe Русский עברית العربيّة فارسی हिंदी বাংলা ภาษาไทย 中文 – 简体 日本語 한국어