MAKED

Large Scale Multi-lingual Keyword Extraction Dataset

Please cite this paper if you use our code or data:

@inproceedings{verma-etal-2022-maked,
    title = "{MAKED}: Multi-lingual Automatic Keyword Extraction Dataset",
    author = "Verma, Yash  and
      Jangra, Anubhav  and
      Saha, Sriparna  and
      Jatowt, Adam  and
      Roy, Dwaipayan",
    editor = "Calzolari, Nicoletta  and
      B{\'e}chet, Fr{\'e}d{\'e}ric  and
      Blache, Philippe  and
      Choukri, Khalid  and
      Cieri, Christopher  and
      Declerck, Thierry  and
      Goggi, Sara  and
      Isahara, Hitoshi  and
      Maegaard, Bente  and
      Mariani, Joseph  and
      Mazo, H{\'e}l{\`e}ne  and
      Odijk, Jan  and
      Piperidis, Stelios",
    booktitle = "Proceedings of the Thirteenth Language Resources and Evaluation Conference",
    month = jun,
    year = "2022",
    address = "Marseille, France",
    publisher = "European Language Resources Association",
    url = "https://aclanthology.org/2022.lrec-1.664/",
    pages = "6170--6179",
    abstract = "Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification. Development and evaluation of keyword extraction techniques require an exhaustive dataset; however, currently, the community lacks large-scale multi-lingual datasets. In this paper, we present MAKED, a large-scale multi-lingual keyword extraction dataset comprising of 540K+ news articles from British Broadcasting Corporation News (BBC News) spanning 20 languages. It is the first keyword extraction dataset for 11 of these 20 languages. The quality of the dataset is examined by experimentation with several baselines. We believe that the proposed dataset will help advance the field of automatic keyword extraction given its size, diversity in terms of languages used, topics covered and time periods as well as its focus on under-studied languages."
}

Please find the Extened - MultiModal MultiLingual Summarization and Keyword Extraction Dataset which is released in a separate publication : HERE

Name		Name	Last commit message	Last commit date
Latest commit History 7 Commits
bengali		bengali
chinese		chinese
english		english
french		french
gujarati		gujarati
hindi		hindi
indonesian		indonesian
japanese		japanese
marathi		marathi
nepali		nepali
pashto		pashto
portuguese		portuguese
punjabi		punjabi
russian		russian
sinhala		sinhala
spanish		spanish
tamil		tamil
telugu		telugu
ukrainian		ukrainian
urdu		urdu
LICENSE		LICENSE
README.md		README.md
fileparser.py		fileparser.py
instructions.md		instructions.md
tokseg.py		tokseg.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

MAKED

About

Uh oh!

Releases

Packages

Languages

License

zenquiorra/MAKED

Folders and files

Latest commit

History

Repository files navigation

MAKED

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages