This dataset of adjective–noun (AN) combinations is extracted from the parsed version of the publicly-available CLC FCE Dataset: https://researchdatasets.cambridge.org/datasets/clc-fce-dataset. The error coding is used to divide the set into two subsets — correctly used ANs and those that are annotated as errors due to inappropriate choice of an adjective or/and noun. For the ANs that are used correctly in some contexts and incorrectly in others, the most frequent annotation from the CLC is used. The dataset contains 4,681 correct and 530 incorrect combinations.
The set of ANs is further divided into corpus-attested and corpus-unattested examples, where a parsed version of the British National Corpus (BNC) is used for reference with the frequency threshold set to 3 occurrences in the corpus.
Both the CLC FCE Dataset and the BNC corpus are lemmatised, tagged and parsed using the RASP system (Briscoe et al., 2006; Andersen et al., 2008).
Further details can be found in the following paper:
Ekaterina Kochmar & Ted Briscoe: ‘Capturing anomalies in the choice of content words in compositional distributional semantic Space’. In Proceedings of Recent Advances in Natural Language Processing (RANLP 2013).
https://aclanthology.org/R13-1047/
Publication date: October 2013
Authors and Contributors
Ekaterina Kochmar, Ted Briscoe
Data security
Please be aware of the problems of leaking benchmark datasets to LLMs (e.g. Balloccu et al, EACL 2024). Please only use this dataset with LLMs hosted locally (e.g. after download from Hugging Face Transformers) or with no retention of data for training if using LLMs via commercial APIs (check that the model cards do not oblige the user to share data or improvements).
Do not provide the corpus (full or partial) to others in any way, even if they have also signed the licence agreement, e.g. through the use of repositories on sites such as Hugging Face and GitHub.
Do not release items (e.g. models, data statistics) derived from the corpus without prior approval of CUP&A.
Citing this paper
Please reference the following paper if you are using this dataset:
@inproceedings{aa2011,
author = {Yannakoudakis, Helen and Briscoe, Ted and Medlock, Ben},
booktitle = {The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies},
title = {A New Dataset and Method for Automatically Grading {ESOL} Texts},
year = {2011}
}
The paper is available at: https://aclanthology.org/P11-1019/.
You may publish the results of research using this dataset. In any such publication you must acknowledge use of the dataset in your research by citing Cambridge University Press & Assessment and the Authors and Contributors as shown.
We ask you to inform us of any such publications by emailing: researchdatasets@cambridge.org
Please report any issues or problems in downloading the dataset by emailing: researchdatasets@cambridge.org
Licence Agreement
- By downloading this dataset and licence, this licence agreement (the “Agreement”) is entered into, effective this date, between you (the “Licensee"), and the Chancellor, Masters and Scholars of the University of Cambridge acting through its department Cambridge University Press & Assessment (the “Licensor”).
- Copyright of the entire licensed dataset is held by the Licensor. No ownership or interest in the dataset is transferred to the Licensee, nor shall the Licensee have any rights in the dataset other than the right to use the dataset in accordance with this Agreement
- The Licensor hereby grants the Licensee a non-exclusive non-transferable right to use the licensed dataset for non-commercial research and educational purposes only. The Licensee shall not sub-licence or assign the benefit or burden of this Agreement in whole or in part.
- Non-commercial purposes exclude without limitation any use of the licensed dataset or information derived from the dataset for or as part of a product or service which is sold, offered for sale, licensed, leased or rented.
- The Licensee shall expressly acknowledge and reference the Licensor when making use of the licensed dataset in all publications of research based on it, in whole or in part, through citation of the paper at the top of the dataset details page.
- The Licensee may publish excerpts of less than 100 words from the licensed dataset pursuant to clause 3.
- The Licensor grants the Licensee this right to use the licensed dataset "as is". Licensor does not make, and expressly disclaims, any express or implied warranties, representations or endorsements of any kind whatsoever. The Licensor has no liability for any loss or damage whatsoever sustained by Licensee as a result of the availability or use of or reliance on the dataset.
- The Licensor shall not be liable for any indirect or consequential loss or damage or for any loss of or corruption of data, loss of programs, profit or goodwill (whether direct or indirect) arising out of or in connection with the access, availability, use of or reliance on the dataset.
- The Licensee shall indemnify and hold the Licensor harmless against any loss or damage which it may suffer or incur as a result of the Licensee’s breach of any terms of this Agreement.
- This Agreement constitutes the entire agreement between the parties and supersedes any previous agreement between the parties relating to its subject-matter. Each party acknowledges and agrees that, in entering into this Agreement, it does not rely on, and shall have no remedy in respect of, any statement, representation, warranty or understanding (whether negligently or innocently made) other than as expressly set out in this Agreement.
- This Agreement shall be governed by and construed in accordance with the laws of England and the English courts shall have exclusive jurisdiction.
You may download this dataset if you agree to the licence terms above and complete the following registration form. Publications using this dataset must acknowledge and reference Cambridge University Press & Assessment as the source of the data.