[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Each row consists of the following fields:
text: The full text, as isannot_text: Annotated text including POS-tagged informationtokens: An ordered list of tokens from the full textpos_tags: Part-of-speech tags for each tokenner_tags: Named entity recognition tags for each tokenNote that by design, the length of tokens, pos_tags, and ner_tags will always be identical.
pos_tags corresponds to the list below:
['SO', 'SS', 'VV', 'XR', 'VCP', 'JC', 'VCN', 'JKB', 'MM', 'SP', 'XSN', 'SL', 'NNP', 'NP', 'EP', 'JKQ', 'IC', 'XSA', 'EC', 'EF', 'SE', 'XPN', 'ETN', 'SH', 'XSV', 'MAG', 'SW', 'ETM', 'JKO', 'NNB', 'MAJ', 'NNG', 'JKV', 'JKC', 'VA', 'NR', 'JKG', 'VX', 'SF', 'JX', 'JKS', 'SN']
ner_tags correspond to the following:
["I", "O", "B_OG", "B_TI", "B_LC", "B_DT", "B_PS"]
The prefix B denotes the first item of a phrase, and an I denotes any non-initial word. In addition, OG represens an organization; TI, time; DT, date, and PS, person.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Thanks to @jaketae for adding this dataset.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Each row consists of the following fields:
text: The full text, as isannot_text: Annotated text including POS-tagged informationtokens: An ordered list of tokens from the full textpos_tags: Part-of-speech tags for each tokenner_tags: Named entity recognition tags for each tokenNote that by design, the length of tokens, pos_tags, and ner_tags will always be identical.
pos_tags corresponds to the list below:
['SO', 'SS', 'VV', 'XR', 'VCP', 'JC', 'VCN', 'JKB', 'MM', 'SP', 'XSN', 'SL', 'NNP', 'NP', 'EP', 'JKQ', 'IC', 'XSA', 'EC', 'EF', 'SE', 'XPN', 'ETN', 'SH', 'XSV', 'MAG', 'SW', 'ETM', 'JKO', 'NNB', 'MAJ', 'NNG', 'JKV', 'JKC', 'VA', 'NR', 'JKG', 'VX', 'SF', 'JX', 'JKS', 'SN']
ner_tags correspond to the following:
["I", "O", "B_OG", "B_TI", "B_LC", "B_DT", "B_PS"]
The prefix B denotes the first item of a phrase, and an I denotes any non-initial word. In addition, OG represens an organization; TI, time; DT, date, and PS, person.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Thanks to @jaketae for adding this dataset.