Orange: Import Documents

From OnnoWiki
Revision as of 08:30, 15 March 2020 by Onnowpurbo (talk | contribs)
Jump to navigation Jump to search

Sumber: https://orange3-text.readthedocs.io/en/latest/widgets/importdocuments.html


Import text document dari folder.

Input

None

Output

Corpus: A collection of documents from the local machine.

Widget Import Documents mengambil file text dari folder dan membuat sebuah corpus. Widget Import Documents dapat membaca .txt, .docx, .odt, .pdf dan .xml. Jika dalam folder ada subfolder, itu dapat digunakan untuk me-label class.

Import-Documents-stamped.png
  • Folder being loaded.
  • Load folder from a local machine.
  • Reload the data.
  • Number of documents retrieved.

If the widget cannot read the file for some reason, the file will be skipped. Files that were successfully retrieved will still be on the output.

Contoh

To retrieve the data, select the folder icon on the right side of the widget. Select the folder you wish to turn into corpus. Once the loading is finished, you will see how many documents the widget retrieved. To inspect them, connect the widget to Corpus Viewer. We’ve used a set of Kennedy’s speeches in a plain text format.

Import-Documents-Example1.png

Now let us try it with subfolders. We have placed Kennedy’s speeches in two folders - pre-1962 and post-1962. If I load the parent folder, these two subfolders will be used as class labels. Check the output of the widget in a Data Table.

Import-Documents-Example2.png


Referensi

Pranala Menarik