Configuring OCR engine
Esta página aún no está disponible en tu idioma.
OpenKM can work with several OCR engines, like Tesseract.
Tesseract
Section titled “Tesseract”Tesseract is an Open Source OCR engine adopted by Goggle. It is an effective option for most use cases. The OCR natively can read TIFF documents and has hight ratio of recognition with images of 300 dpi of resolution and converted to lineart (1 bit color).
The supported image formats are:
- TIFF
- PNG
- JPG
- GIF
Installation
Section titled “Installation”| OS | Description |
|---|---|
| Ubuntu and Debian | bash<br>$ apt-get install tesseract-ocr tesseract-ocr-eng<br>Depending on your locale you may like to install other language files. More information at: - https://github.com/tesseract-ocr - https://code.google.com/p/tesseract-ocr/ |
| Red Hat and CentOS | More information at: - https://code.google.com/p/tesseract-ocr/wiki/Compiling - https://code.google.com/p/python-tesseract/wiki/HowToCompileForCentos - http://www.vicchiam.com/blog/?p=168 |
| Windows | Download from https://code.google.com/p/tesseract-ocr/ and follow the installation wizard. |
Check installation
Section titled “Check installation”From the command line execute (where “image.jpg” should be an existing image file):
$ tesseract image.jpg textConfiguration
Section titled “Configuration”Go to Administration > Configuration parameters and edit the parameter named system.ocr:
| OS | |
|---|---|
| Linux | /usr/bin/tesseract ${fileIn} ${fileOut} or /usr/bin/tesseract ${fileIn} ${fileOut} -l esp The parameter -1 esp must be used for the Spanish language support while executing OCR. To see all tesseract options execute: bash<br>$ tesseract<br> |
| Windows | c:\Tesseract-ocr-3.0.2\tesseract.exe ${fileIn} ${fileOut} or c:\Tesseract-ocr-3.0.2\tesseract.exe ${fileIn} ${fileOut} -l spa The parameter -1 esp must be used for the Spanish language support while executing OCR. To see all tesseract options execute: bash<br>c:\Tesseract-ocr-3.0.2> tesseract.exe<br> |
- Parameter ${fileIn} will be replaced internally by the image document to be processed.
- Parameter ${fileOut} will be replaced internally by the temporal text file.
- You can set other additional parameters like the -l spa parameter, for example.
Additional information
Section titled “Additional information”- http://code.google.com/p/tesseract-ocr/
- Tesseract - Summary & first experiences.
- Tesseract OCR Google Groups
- First Interactions with Tesseract OCR on Ubuntu Linux
- http://code.google.com/p/tesseract-ocr/wiki/ReadMe
- http://code.google.com/p/tesseract-ocr/wiki/FAQ
Dictionary
Section titled “Dictionary”You can also use an OpenOffice.org dictionary to enhance the OCR process. You can find these language specific dictionaries at OpenOffice.org Dictionary Repository.
Configuration
Section titled “Configuration”Go to Administration > Configuration parameters and edit the parameter named system.openoffice.dictionary:
| OS | |
|---|---|
| Linux | /home/openkm/dictionary/en-GB.zip |
| Windows | c:\tomcat-7.0.61\dictionary\en-GB.zip |
OCR rotate configuration
Section titled “OCR rotate configuration”There’s a configuration parameter that forces document ocr on rotated documents.
Configuration
Section titled “Configuration”Go to Administration > Configuration parameters and edit the parameter named system.ocr.rotate:
| Property | Description |
|---|---|
| system.ocr.rotate | The parameter is a collection of degrees separated by character “;” 0;90;180;270; |