Skip to main content

Watched Folders

The Watched folders system template lets NomaUBL fetch source documents straight from directories — typically the network share where JD Edwards drops its report outputs — instead of waiting for files to be pushed into the input folder. Each run lists the configured folders, keeps the new files, converts PDFs to XML in parallel (XML files are copied as is) into the document template's input folder, then runs one batch per template.

Open it under Configuration → System → Watched folders. The template ships empty: nothing runs until a folder is configured and a job is scheduled.


At a glance​

Watched folder/data/mnt/jde_outputspattern R42565_*new since last scan − overlapalready processed → skippedConversionPDF → XML (pdf2xml)in parallelXML copied as isTemplate inputdirInput/invoicesBatchone per templateUBL · PDF · PA

Folders​

Each row of the table is one watched folder; + Add folder appends a row, × removes it.

ColumnDescription
DirectoryThe folder to scan (e.g. /data/mnt/jde_outputs).
PatternFile-name pattern of the files to pick up (e.g. R42565_*).
TemplateThe document template the files are processed with — they land in its input folder.
ManifestOptional pdf2xml manifest used to convert the PDFs of this folder.
Overlap hOverlap window in hours (default 2): each run looks at files newer than the last scan minus this window.
MountWhen ticked, the directory must be a mounted share — an unmounted mount point is reported as an error instead of being read as empty.
Last scanThe watermark written by every run (UTC instant). Clear it or type a date (yyyy-MM-dd, yyyy-MM-ddTHH:mm:ss in server time, or an ISO instant) to rescan from that point.

Above the table, Parallel conversions (default 4) sets how many PDFs are converted at once; the batch itself keeps its own thread pool.


How a run works​

  1. List — every folder is listed for files matching its pattern and newer than last scan − overlap. A folder scanned for the first time, without a date, only takes the overlap window — use a backfill date for history.
  2. Filter — files already processed (source file name found in the archive) or already waiting as XML in the template's input folder are dropped.
  3. Convert — PDFs are converted to XML in parallel into the template's input folder; XML files are copied as is.
  4. Process — one batch runs per template, with the batch's own parallelism.

The last-scan date advances on every run. A file that fails is retried while it stays inside the overlap window and reported by name — it never blocks the others.


Run now​

The Run now group runs the saved configuration on demand:

  • Dry run — only lists what would be converted and processed.
  • Scan and process — runs the full pickup.
  • An optional backfill date widens every folder's window for this run only.

Scheduling and command line​

  • Scheduled — under global → Scheduling → Batch Document Processing, add a batch job with the source Watched folders, every N minutes or daily at a fixed time. The job can target a single folder (picked from the configured directories) or all of them.
  • Command line — nomaubl.sh watch-folders <env> [--folder N] [--since <date>] [--dry-run] [--no-process].
  • API — POST /api/watch-folders/run.

Tips & best practices​

  • Keep the overlap larger than the copy time. A file still being written when the scan passes is picked up on the next run as long as it stays inside the window.
  • Tick Mount on network shares. An unmounted share otherwise looks like an empty folder and nothing is processed, silently.
  • Start with a dry run. It shows exactly which files a new folder row would pick up before anything is converted.