Building a dual dataset of text- and image-grounded conversations and summarisation in Gà idhlig (Scottish Gaelic)
Name
Building_a_dual_dataset_of_text_and_image_grounded_conversations.pdf
Description
visibility:open
Size
1.1 MB
Format
Adobe PDF
Checksum (CRC64NVME)
mHv/s1TTr+M=
Resource type
Book chapter
Creator (person)
Date published
2023
Abstract
Gà idhlig (Scottish Gaelic; gd) is spoken by about 57k people in Scotland,1 but remains an under-resourced language with respect to natural language processing in general and natural language generation (NLG) in particular. To
address this gap, we developed the first datasets for Scottish Gaelic NLG, collecting both conversational and summarisation data in a single setting. Our task setup involves dialogues between a pair of proficient speakers discussing
museum exhibits, grounding the conversation in images and texts. Then, each interlocutor summarises the dialogue resulting in a secondary dialogue summarisation dataset. This paper presents the dialogue and summarisation
corpora, as well as the software used for data collection. The dialogue dataset consists of 43 conversations (13.7k words) and 61 summaries (2.0k words).2
Project(s)
NLG for low-resource domains
Funder
| Funder name | Awards |
EPSRC Centre for Doctoral Training in Technology Enhanced Chemical Synthesis | EP/T024917/1 |
Book title
The 16th International Natural Language Generation Conference Proceedings of the Conference September 11 - 15, 2023
Pagination
443-448
Publisher
The Association for Computational Linguistics
Related URL
Rights statement
In Copyright
Additional information
Full paper available via the official URL