Difference between revisions of "Resources for Bulgarian"

From ACL Wiki
Jump to: navigation, search
(Machine translation systems)
(6 intermediate revisions by 3 users not shown)
Line 2: Line 2:
  
 
===Free software===
 
===Free software===
 +
 +
* [https://apertium.svn.sourceforge.net/svnroot/apertium/trunk/apertium-mk-bg apertium-mk-bg] RBMT system between Macedonian and Bulgarian
  
 
===Proprietary===
 
===Proprietary===
Line 8: Line 10:
  
 
==Lexical resources==
 
==Lexical resources==
 +
===Morphological analysis===
 +
 +
====Free software====
 +
 +
* [https://apertium.svn.sourceforge.net/svnroot/apertium/trunk/apertium-mk-bg/apertium-mk-bg.bg.dix Morphological analyser] 8,581 lemmata, ~88% coverage over SETimes
 +
 +
====Proprietary====
 +
 +
== Grammars ==
  
 
===Proprietary===
 
===Proprietary===
  
 
* [http://dcl.bas.bg/BulNet/general_en.html BulNet WordNet] (21,444 synonym sets)
 
* [http://dcl.bas.bg/BulNet/general_en.html BulNet WordNet] (21,444 synonym sets)
 +
* [[Generation grammars|KPML generation grammar]]
  
 
==Corpora==
 
==Corpora==
  
 
===Free===
 
===Free===
* [http://www.hf.uio.no/easteur-orient/bulg/mat/ Corpus of spoken Bulgarian]
+
 
* [http://xixona.dlsi.ua.es/~fran/setimes/ Southeast European Times] (paragraph aligned corpus, Albanian,Bulgarian,English,Greek,Macedonian,Romanian,Serbo-Croatian,Turkish — 9,678 paragraphs, 92,450— 122,912 words per language)
+
* [http://www.statmt.org/setimes/ Southeast European Times] (sentence aligned corpus, Albanian, Bulgarian, English, Greek, Macedonian, Romanian, Serbo-Croatian, Turkish — approximately 4.5 million words per language)
  
 
===Proprietary===
 
===Proprietary===
 +
 +
* [http://www.hf.uio.no/easteur-orient/bulg/mat/ Corpus of spoken Bulgarian]
  
 
==Bibliography==
 
==Bibliography==
  
*
 
  
 
==External links==
 
==External links==
  
*
 
  
 
[[Category:Resources by language|Bulgarian]]
 
[[Category:Resources by language|Bulgarian]]

Revision as of 17:04, 7 October 2010

Machine translation systems

Free software

Proprietary

Lexical resources

Morphological analysis

Free software

Proprietary

Grammars

Proprietary

Corpora

Free

  • Southeast European Times (sentence aligned corpus, Albanian, Bulgarian, English, Greek, Macedonian, Romanian, Serbo-Croatian, Turkish — approximately 4.5 million words per language)

Proprietary

Bibliography

External links