Document - Data Source

Use one or more documents as a data source. This section covers how to upload documents.

Click on Document to add documents as a data source:

Add data source menu

Upload

Upload your files (click or drag and drop on the dedicated area):

Info
The following file formats are currently supported: .txt, .html, .md, .ods, .docx, .xlsx, .doc, .rtf, .odt, .csv, .pdf, .pptx, and image files (.png, .jpg, .jpeg, .gif, .bmp, .webp, .tiff).
document empty

Extracting text from images and scans

You can upload image files (.png, .jpg, .gif, .bmp, .webp, .tiff) and scanned PDFs as a data source, just like any other document. By default only the text that already exists in the file is indexed — text that lives inside an image is not read. To make that text searchable, turn on one of the two extraction modes below for the file before indexing.

Each row in the file list has a toggle on its right-hand side. OCR and LLM are mutually exclusive, so enabling one automatically turns the other off.

Optical Character Recognition (OCR)

OCR converts a text image into machine-readable text. For example, scanning a document produces an image file that cannot be edited, searched, or word-counted directly.

OCR converts that image into a text document whose contents are stored as text data.

Warning
OCR results may not be 100% accurate. Review the extracted text for errors, especially with handwriting or low-quality scans.

LLM parsing

LLM parsing sends the image (or scanned page) to a vision-capable language model that reads it the way a person would. It generally handles complex layouts, tables, and handwriting better than OCR, and can also describe non-text visual content.

Info
LLM parsing is only available to administrators and to users who have the "LLM File Parsing" entitlement enabled. Running it consumes part of your LLM quota.

To enable a mode, click the OCR or LLM switch on the right-hand side of the file you want to apply it to:

documents list OCR

Once all documents are specified, click "Finish". The data source page opens:

documents list
Info
Uploaded images are converted to a one-page PDF, so they can be previewed and downloaded from the data source just like any other document.

Batch Actions on Selected Files

You can select multiple files at once using the checkboxes on the left side of each row.

Once one or more files are selected, a batch action toolbar appears above the table with the following actions:

  • Downloaddownloads all selected PDF files at once. The button is active only when at least one selected file is a PDF or audio file in Indexed status.
  • Deletepermanently removes all selected files from the connector.
Document table with batch download and delete toolbar
Tip

To try out this data source, use this collection of PDFs about the Simpsons:

These documents are coming from the Simpsons wiki

Access a live search interface built on this PDF collection: