<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
  <teiHeader type="ESBD-ER">
      <fileDesc>
         <titleStmt>
            <title type="main">A Spoken Corpus of Cameroon Pidgin English: pilot study</title>
            <author>Dr. Melanie Green, University of Sussex</author>
	           <author>Dr. Miriam Ayafor, University of Yaoundé I</author>
	           <author>Dr. Gabriel Ozon, University of Sheffield</author>
	           <funder>British Academy / Leverhulme Trust</funder>
         </titleStmt>
         <editionStmt>
            <edition/>
         </editionStmt>
         <extent>
            <seg type="designation">CollectionSound</seg>
            <seg type="size">264 files: ca. 2.6 GB</seg>
         </extent>
         <publicationStmt>
            <authority>deposited by<name type="person">Melanie Green</name>
               <name type="institution">University of Sussex</name>
               <date>2016-09-14</date>
            </authority>
            <distributor>
               <orgName type="OTA">Oxford Text Archive</orgName>
               <placeName>Oxford</placeName>
               <address>
                  <orgName type="OTA">Oxford Text Archive</orgName>
                  <orgName type="Bodleian">Bodleian Libraries</orgName>
                  <addrLine>Broad St</addrLine>
                  <addrLine>Oxford</addrLine>
                  <addrLine>OX1 3BG</addrLine>
               </address>
               <email>ota@bodleian.ox.ac.uk</email>
            </distributor>
            <idno type="OTA">2563</idno>
            <idno type="ota">2563</idno>
	           <idno type="PURL">http://purl.ox.ac.uk/ota/2563</idno>
            <availability n="OTA-new" status="free" rend="visible">
               <licence target="http://creativecommons.org/licenses/by-nc-sa/3.0/">
          Distributed by the University of Oxford under a Creative Commons Attribution-NonCommercial-ShareAlike 3.0 Unported License.
          </licence>
            </availability>
            <date>2016-09-14</date>
         </publicationStmt>
         <notesStmt>
 	          <note>Publications based on the data include:
		<list>
                  <item>Ayafor, Miriam and Melanie Green (2017). Cameroon Pidgin English: A comprehensive grammar [London Oriental and African Language Library 20]. Amsterdam: John Benjamins.</item>
		                <item>Ozón, Gabriel, Melanie Green, Miriam Ayafor and Sarah FitzGerald (2017). Building a spoken corpus of Cameroon Pidgin English: methodological challenges. World Englishes. ISSN 0883-2919. </item>
		                <item>Green, Melanie and Gabriel Ozón (2017). Information structure in a spoken corpus of Cameroon Pidgin English. In Adamou, Evangelia; Haude, Katharina and Vanhove, Martine (eds.) Information structure in lesser-described languages: Studies in prosody and syntax. Amsterdam: John Benjamins.</item>
		                <item>Green, Melanie and Gabriel Ozón (2015). Valency and transitivity in contact: evidence from Cameroon Pidgin English. Journal of Language Contact. ISSN 1877-4091.</item>
		             </list>
	           </note>
            <note>
		             <p>The POS-tagging was carried out by Sarah Fitzgerald.</p>
	           </note>
         </notesStmt>
         <sourceDesc>
            <bibl>        
               <p>This resource is a 240,000-word corpus of spoken Cameroon Pidgin English (CPE), a widely-used yet stigmatised and largely uncodified pidgin/creole variety.</p>
               <p>The corpus consists of transcriptions of private and public dialogues and monologues, with mark-up and POS-tagging, together with accompanying sound files. The recordings were conducted in five different locations in Cameroon (Bamenda, Buea, Douala, Kumba and Yaounde), allowing some insights into regional variation. Text categories and the proportions of monologue and dialogue are guided by those of the International Corpus of English (ICE) project, which makes the corpus immediately comparable with existing corpora of post-colonial varieties of English.</p>
               <p>
                  <list>
                     <item>Spelling: since there is no standardised orthography for CPE, the orthography adopted for this project is based on that developed by Ayafor (2014), which was kept under review during the course of the project.</item>
                     <item>
Annotation was added to the transcriptions based on ICE guidelines for the annotation of spoken texts: standard mark-up symbols were used to denote text unit, speaker identification, overlapping speech, unclear words, uncertain transcriptions, anthropo-phonics, editorial comments, foreign words and indigenous language words.</item>
                     <item>
Tagging: a tagset for CPE was devised based on CLAWS 5. Initially tagging was conducted manually, and then by means of TreeTagger. A third of the corpus has been post-checked, with accuracy rates at 94%. </item>
                  </list>
               </p>
               <p>
The corpus is aimed at providing a resource for linguistic description and comparison. It allows linguists to identify and describe recurring grammatical patterns, as well as the phonology of the language (given the availability of sound files deposited with the text files). It also allows comparison of CPE with other pidgin/creole languages, other Cameroonian and West African languages, and other varieties of post-colonial English. Furthermore, the corpus provides an exceptional resource for the study of general/theoretical linguistics, creolistics, typology, language contact and change, sociolinguistics and discourse analysis.</p>
               <p>The corpus contains 80 sound recordings of monologues (scripted and unscripted) and dialogues (public and private). Each sound file (in .wav format) is 10-15 minutes in length. These recordings have been transcribed (each approximately 3,000 words in length) and annotated. Transcriptions are submitted in two formats: (a) plain transcription (with basic markup indicating speaker turns, overlaps, etc.), and (b) a POS-tagged version, which adds POS-tags to the plain version of the transcription.</p>
               <p>The language of the monologues is Cameroon Pidgin English, with codeswitching into English, French, and indigenous Cameroonian languages.</p>
               <p>
The accompanying documentation includes (i) a list of submitted files, (ii) a list of participant data, (iii) a tagging guide, (iv) a word list and spelling guide.</p>
		          </bibl>
         </sourceDesc>
      </fileDesc>
      <encodingDesc>
         <projectDesc>
            <p>This resource is a 240,000-word corpus of spoken Cameroon Pidgin English (CPE), a widely-used yet stigmatised and largely uncodified pidgin/creole variety. 
</p>
            <p>
The corpus consists of transcriptions of private and public dialogues and monologues, with mark-up and POS-tagging, together with accompanying sound files. The recordings were conducted in five different locations in Cameroon (Bamenda, Buea, Douala, Kumba and Yaounde), allowing some insights into regional variation. Text categories and the proportions of monologue and dialogue are guided by those of the International Corpus of English (ICE) project, which makes the corpus immediately comparable with existing corpora of post-colonial varieties of English. 
</p>
            <p>
               <list>
                  <item>Spelling: since there is no standardised orthography for CPE, the orthography adopted for this project is based on that developed by Ayafor (2014), which was kept under review during the course of the project.</item>
                  <item>
Annotation was added to the transcriptions based on ICE guidelines for the annotation of spoken texts: standard mark-up symbols were used to denote text unit, speaker identification, overlapping speech, unclear words, uncertain transcriptions, anthropo-phonics, editorial comments, foreign words and indigenous language words.</item>
                  <item>
Tagging: a tagset for CPE was devised based on CLAWS 5. Initially tagging was conducted manually, and then by means of TreeTagger. A third of the corpus has been post-checked, with accuracy rates at 94%. </item>
               </list>
            </p>
            <p>
The corpus is aimed at providing a resource for linguistic description and comparison. It allows linguists to identify and describe recurring grammatical patterns, as well as the phonology of the language (given the availability of sound files deposited with the text files). It also allows comparison of CPE with other pidgin/creole languages, other Cameroonian and West African languages, and other varieties of post-colonial English. Furthermore, the corpus provides an exceptional resource for the study of general/theoretical linguistics, creolistics, typology, language contact and change, sociolinguistics and discourse analysis.</p>
            <p>The corpus was created between September 2014 and August 2016.</p>
         </projectDesc>
         <editorialDecl>
            <p>Encoding format: audio .wav files (RIFF little-endian, Microsoft PCM, 44.1 KHz, 16-bit stereo); text transcripts .txt files; POS-tagged text transcripts in .txt files</p>
         </editorialDecl>
         <classDecl>
            <taxonomy xml:id="OTASH">
               <bibl>University of Oxford Text Archive Subject Headings</bibl>
            </taxonomy>
            <taxonomy xml:id="LCSH">
               <bibl>Library of Congress Subject Headings</bibl>
            </taxonomy>
         </classDecl>
      </encodingDesc>
      <profileDesc>
         <creation>
            <date>2014-2016</date>
         </creation>
         <langUsage>
            <language ident="wes">Cameroon Pidgin</language>
         </langUsage>
         <textClass>
            <keywords scheme="#OTASH">
               <term type="genre">Linguistic corpora</term>
            </keywords>
            <keywords scheme="#OTASH">
               <term type="resource_type">Corpus</term>
            </keywords>
            <keywords scheme="#OTASH">
	              <term type="genre">Speech--Research</term>
            </keywords>
            <keywords scheme="#LCSH">
               <term type="genre">Linguistics</term>
            </keywords>
            <keywords scheme="#LCSH">
               <term type="genre">Linguistics analysis (Linguistics)</term>
            </keywords>
            <keywords scheme="#LCSH">
               <term type="genre">Speech--Synthesis</term>
            </keywords>
            <keywords scheme="#LCSH">
               <term type="genre">Pidgin English</term>
            </keywords>
            <keywords scheme="#LCSH">
               <term type="genre">Code switching (Linguistics)</term>
            </keywords>
         </textClass>
         <textClass>
            <keywords scheme="http://www.ota.ox.ac.uk/processing">
               <term type="class">corpus</term>
               <term type="license">Free for non-commercial use</term>
            </keywords>
         </textClass>
      </profileDesc>
      <revisionDesc>
         <change>
            <date>2016-09-14</date>
            <label>cataloguer</label>
            <name>Wynne, Martin</name>TEIXML header composed
 </change>
      </revisionDesc>
  </teiHeader>
  <text>
      <body>
         <div type="OTA-filemap" rend="visible" n="unrestricted">
            <ab type="dir">
               <seg type="part">1</seg>
               <ptr type="resourcePart" target="Cameroon-Pidgin-English"/>
            </ab>
         </div>
      </body>
  </text>
</TEI>
