RIRE corpus

View Dataset
Bories, Anne-Sophie;Nugues, Lara;Couturier, Nils;Plechac, Petr

Description

The JIGS (Joke-like IncongruityGathering System) was developed as part of the SNSF-funded project “Le Rire des vers / Mining the comic verse”. It is used to tag JOLIS (Joke-Like Incongruity Segments)A JOLIS (joke-like incongruity segment) is a text segment where:Two possible meanings are present (S1/S2), andThe two meanings (scripts) are ovelapping in the same segment, andThe two scripts are incompatible.If all three criteria are met, the segment is a JOLI-SThe corpus featured in the “Le Rire des vers / Mining the comic verse” contains 8255 French poems that have been annotated for versification features by means of Richard Renault's Malherbe. Some of these poems have been obtained from the anamètre database, others have been gathered and prepared by us. Part of the poems has been annotated for joke-like incongruities and for tunes they are associated with (tune annotation is still preliminary).Corpus is stored in a single JSON file and simultaneously in XML files (one XML file per poem).XML structureThe structure of each XML files is as follows:

[title of the poem] [author of the poem] [title of the book the poem comes from] Le Rire des vers
[text of the segment]... ... ... ... ...
The attributes of the elements are:

id: id of the poemtune_name: name of the associated tune as stated in the text (optional)tune_name_standardized: standardized name of the tune (optional)tune_genre: genre of the tune (optional)tune_composer: composer of the tune (optional)for other attributes see https://crisco4.unicaen.fr/verlaine/index.php?navigation=descriptionid: id of the stanzatype: stanza type, e.g. quatrainrhyme: rhyme scheme, e.g. abbaid: id of the linetext: text of the linefor other attributes see https://crisco4.unicaen.fr/verlaine/index.php?navigation=descriptionid: id of the wordtext: text of the wordpos_stanford: part-of-speech tag assigned Stanford Taggerpos_treetagger: part-of-speech tag assigned by TreeTaggerlemma_treetagger: part-of-speech tag assigned by TreeTaggerid: id of the segmentfor other attributes see https://crisco4.unicaen.fr/verlaine/index.php?navigation=descriptionid: unique id of the jolifrom: id of word element where joli startsto: id of word element where joli endsso_actual_non: [Only one script is real/actualised], or [Only one script is literal (vs figurative)] so_normal_abnormal: [Only one script is normal]so_possible_impossible: [Only one script is possible]so_good_bad: [Only one script is positive or negative]so_life_death: [Script pitting life against death], or [Script pitting animated against inanimate]so_obscenity: [Sexual/satological reference]so_money: [Mention of money]so_high_low_stature: [Mention of high/low stature]so_human_non: [Only one script is human]s1_s2: open field: [Short description of script 1 and script 2]si: open field: [Brief description of the situation]ta: open field: short mention of the target (if any)la: open field: remarks about the languagela_argot: [use of slang]la_cacography: [use of misspellings]comment: optional commentannotator: id of the annotator/from: id of word element where jabline/punchline startsto: id of word element where jabline/punchline endscomment: optional commentJSON structureJSON file follows similar logic as the XMLs. It is (in Python terms) a list of dicts representing individual poems. Beside keys listed above under
attributes, each poem holds an “lg” key, which contains a list of dicts representing individual stanzas. Beside keys listed above under attributes, each stanza holds an “l” key, which contains a list of dicts representing individual lines…JOLIs are stored under the “jolis” key of individual poems.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.5

FAIR Score

81%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Zenodo

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Artificial Intelligence

Field

Computer Science

Domain

Physical Sciences

Confidence Score

37%

Source

Scholar Data Model

Normalization Factors

FT

57.69

CTw

1.00

MTw

1.00