Web - Amazon

We provide Linux to the World


We support WINRAR [What is this] - [Download .exe file(s) for Windows]

CLASSICISTRANIERI HOME PAGE - YOUTUBE CHANNEL
SITEMAP
Audiobooks by Valerio Di Stefano: Single Download - Complete Download [TAR] [WIM] [ZIP] [RAR] - Alphabetical Download  [TAR] [WIM] [ZIP] [RAR] - Download Instructions

Make a donation: IBAN: IT36M0708677020000000008016 - BIC/SWIFT:  ICRAITRRU60 - VALERIO DI STEFANO or
Privacy Policy Cookie Policy Terms and Conditions
Document classification - Wikipedia, the free encyclopedia

Document classification

From Wikipedia, the free encyclopedia

Document classification/categorization is a problem in information science. The task is to assign an electronic document to one or more categories, based on its contents. Document classification tasks can be divided into two sorts: supervised document classification where some external mechanism (such as human feedback) provides information on the correct classification for documents, and unsupervised document classification, where the classification must be done entirely without reference to external information.

Contents

[edit] Techniques

Document classification techniques include:

and approaches based on natural language processing.

[edit] Applications

A recent notable use of document classification techniques has been spam filtering which tries to discern E-mail spam messages from legitimate emails.

[edit] See also

[edit] External links

Publications:

Resources:

Data sets:

Software:

  • LingPipe - Java natural language processing software including a rich classification runtime and evaluation framework with classifiers based on character- and token- language models (including Naive Bayes).
  • TIS eFLOW platform - a modular solution that offers advanced data capture and document classification capabilities.
  • YALE (Yet Another Learning Environment) - freely available integrated open-source software environment for knowledge discovery, data mining, machine learning, visualization (e.g. of text clusterings), etc. featuring a plugin WordVectorTool for text mining tasks like text classification, text clustering, document feature set construction and transformation, etc.
  • Bow - freely available open-source toolkit for statistical language modeling, text retrieval, classification, and clustering.
  • XmlMiner Data and text mining toolkit targeted at XML data.
Our "Network":

Project Gutenberg
https://gutenberg.classicistranieri.com

Encyclopaedia Britannica 1911
https://encyclopaediabritannica.classicistranieri.com

Librivox Audiobooks
https://librivox.classicistranieri.com

Linux Distributions
https://old.classicistranieri.com

Magnatune (MP3 Music)
https://magnatune.classicistranieri.com

Static Wikipedia (June 2008)
https://wikipedia.classicistranieri.com

Static Wikipedia (March 2008)
https://wikipedia2007.classicistranieri.com/mar2008/

Static Wikipedia (2007)
https://wikipedia2007.classicistranieri.com

Static Wikipedia (2006)
https://wikipedia2006.classicistranieri.com

Liber Liber
https://liberliber.classicistranieri.com

ZIM Files for Kiwix
https://zim.classicistranieri.com


Other Websites:

Bach - Goldberg Variations
https://www.goldbergvariations.org

Lazarillo de Tormes
https://www.lazarillodetormes.org

Madame Bovary
https://www.madamebovary.org

Il Fu Mattia Pascal
https://www.mattiapascal.it

The Voice in the Desert
https://www.thevoiceinthedesert.org

Confessione d'un amore fascista
https://www.amorefascista.it

Malinverno
https://www.malinverno.org

Debito formativo
https://www.debitoformativo.it

Adina Spire
https://www.adinaspire.com