Difference between revisions of "Resources for Arabic"

From ACL Wiki
Jump to: navigation, search
(Corpora)
m (Free/open licence: quranic arabic corpus)
Line 14: Line 14:
 
===Free/open licence===
 
===Free/open licence===
 
* [http://github.com/anastaw/Meedan-Memory Meedan-Memory], Arabic-English TMX (sentence-aligned), ~467,000 words on the English side, [http://www.opendatacommons.org/licenses/odbl/ Open Database Licence]
 
* [http://github.com/anastaw/Meedan-Memory Meedan-Memory], Arabic-English TMX (sentence-aligned), ~467,000 words on the English side, [http://www.opendatacommons.org/licenses/odbl/ Open Database Licence]
 +
* [http://quran.uk.net/ Quranic Arabic Corpus], 77,430 words of Quranic Arabic, with manually verified contextual POS, inflection, derivation; [[dependency grammar]] annotation is planned.
  
 
==Parser==
 
==Parser==

Revision as of 11:22, 12 November 2009

Morphology

Free software

  • AraMorph - Perl - An Arabic morphological analyzer and part-of-speech tagger written in Perl (originally by Tim Buckwalter)
  • AraMorph - Java - An Arabic morphological analyzer and part-of-speech tagger rewritten in Java for Lucene

Proprietary

Corpora

Proprietary

Free/open licence

Parser

Free software

Bibliography

External links