Files
FCCS/README.md

9.4 KiB

FCCS Data Extraction Tool

Extracts and organizes documents from FileCabinet CS (Thomson Reuters) using GUI automation. Designed to work across multiple client engagements without code changes.

Prerequisites

  • 32-bit Python (must match FCCS architecture)
  • pywinauto (pip install pywinauto)
  • Pillow (pip install Pillow) — required for failure screenshots
  • FileCabinet CS installed and restored with client backup data

Project Layout

File Purpose
config.template.ini Tracked template for all paths, timeouts, and FCCS control identifiers
config.ini Per-machine working config (git-ignored; auto-created from the template)
fccs_folders.txt FCCS folder templates (manually populated per engagement)
fccs_config.py Shared config loading, logging, folder pattern building
fccs_scan.py Step 1: Scan backup directory for drawer IDs
fccs_export.py Step 2: Automate FCCS GUI to export all drawers
fccs_reorganize.py Step 3: Parse filenames and rebuild folder structure
fccs_verify.py Step 4 (optional): Compare manifests against exported files
fccs_check.py Utility: Interactively check if specific drawers exported completely
fccs_report.py Utility: Generate a clean client-facing HTML progress report (issues only)
fccs_dump_controls.py Utility: Dump control identifiers of an on-screen dialog

Setup Per Engagement

  1. Edit config.ini -- point backup_dir, export_dir, and output_dir at the client's directories. This file is created automatically from config.template.ini the first time you run any script (it's git-ignored, so your machine-specific paths never drift the repo). To reset a machine, delete config.ini and re-run. Keep shared/default changes in config.template.ini (the tracked file).
  2. Populate fccs_folders.txt -- open FCCS > System Configuration > Document Folders and list each folder template exactly as shown, one per line. Keep the YYYY prefix on recurring folders. UltraTax CS folders are detected automatically and do not need to be listed.

Workflow

Step 1: Scan Drawers

python fccs_scan.py

Reads the restored FCCS backup directory and writes all drawer IDs (subfolder names) to drawer_ids.txt. It also:

  • Flags prefix clashes — if one drawer ID is a prefix of another (e.g. 02218 and 02218A), searching the base ID in FCCS pops up a selection box that breaks plain automated navigation. These clashes are reported, and the base (shorter) IDs are auto-seeded into ignore.txt. The longer, more-specific IDs (02218A) search fine and export normally; the base IDs are handled by the separate clash-export script.
  • Reports ignored drawers — any IDs listed in ignore.txt that exist in this backup are shown as ones the export will skip.

Ignoring drawers: The scan creates ignore.txt (at ignore_file, default C:\Migration\ignore.txt) if it doesn't exist and pre-fills it with the clash base IDs — searching those in FCCS shows a selection box that stalls the plain export, so they're skipped by the main export. Open the file and:

  • Delete or comment out any clash base you'd rather handle fully by hand.
  • Add any other drawers to skip (e.g. password-protected folders), one ID per line.

Re-running the scan never overwrites your edits — it only appends newly-discovered clashes. drawer_ids.txt stays a full inventory; the export skips anything active in the ignore list. Lines starting with # are comments.

Step 2: Export Documents

python fccs_export.py

Requires FCCS to be open with export destination already configured. Automates the GUI to export every drawer via File > Send To > File. Features:

  • Resumable -- tracks completed drawers in completed.txt; safe to restart
  • Screenshots -- captures failure states for diagnosis
  • Defensive -- one bad drawer won't crash the entire run
  • Crash recovery -- some documents (e.g. UltraTax "Diagnostics" files) crash FCCS's converter (FileConversionEngine::convert() failed), which aborts that drawer's export. The script detects the error dialog, screenshots and logs the crashing document, dismisses it, and records the drawer in crashed.txt so it's skipped on future runs instead of stalling. Handle crashed drawers manually (export them excluding the poison document); delete a line from crashed.txt to retry after fixing.

Step 3: Reorganize Files

python fccs_reorganize.py

Parses the flat exported filenames and copies them into an organized structure:

output/
  ABRAHAM, REBEKAH L./
    Tax Documents/
      2025/
        030126 E-mail re Tax Info.pdf
    Billing & Invoices/
      2026/
        030826 Invoice for 2025 Forms 1040 & IL-1040.pdf
    UltraTax CS/
      12-31-2008/
        2008 Form 1040 Filing Instructions.doc
    Permanent File/
      Driver's License.pdf
    _unparsed/
      (files for this client that couldn't be fully parsed, kept for review)
  _unparsed/
    (only files whose client couldn't be recovered from the filename)

Exported filenames follow the format {drawer_id}_{client_name}_{folder_name}_{creation_date}_{document_name}.ext. The parser uses folder templates from fccs_folders.txt (with YYYY expanded via regex) and the creation date (MM-DD-YYYY) as anchors to reliably split the underscore-delimited fields. Folder names are decomposed into nested paths that match the FCCS UI structure (e.g. 2025 Tax Documents becomes Tax Documents/2025/). UltraTax CS folders are matched by a built-in pattern.

Files that can't be fully parsed still keep their client: the drawer ID and client name are the first two underscore-delimited tokens and stay recoverable even when the folder/date parse fails, so those files are filed under {client_name}/_unparsed/ (retaining their original filename). Only files whose client can't be recovered at all fall back to the top-level _unparsed/.

Folder names are sanitized for Windows before use — trailing spaces and periods are stripped from each path component, since Windows can't create a directory ending in a space or dot (e.g. a client named FARR GROUP, P.L. becomes FARR GROUP, P.L). Reorganization is also resilient per file: if one file can't be placed for any reason, the error is logged and counted (reported as errored in the summary) and the run continues with the rest rather than aborting.

Step 4 (optional): Verify Export Completeness

During export, each drawer's document list is captured from the FCCS dialog and saved as a manifest (in manifest_dir). These tools compare the manifests against the files actually in the export folder to confirm nothing was missed.

Both compare at the document level and share identical matching logic. They account for:

  • Page-splitting -- a multi-page document exported as Name Page 1, Name Page 2, … counts as that one document being present.
  • Filename sanitization -- document titles containing characters illegal in Windows filenames (e.g. : / ?) still match the exported files.

Batch-check every drawer that has a manifest:

python fccs_verify.py

Reports each drawer as OK or INCOMPLETE (listing the missing documents), plus a summary and any exported drawers that have no manifest. The full report is written to its own file (verify_report, default C:\Migration\verify_report.txt) as well as the console — separate from the export's run_log.txt.

Spot-check specific drawers interactively (e.g. ones the log marked failed, to see whether they actually finished exporting in the background):

python fccs_check.py
Drawer ID(s): 08097 18430

Note: because page-splitting means the number of files can't be mapped one-to-one to documents, completeness is judged by document presence (is each manifest document represented by at least one exported file), not by exact file counts.

Client-facing progress report: For a clean summary to share with the client, generate an HTML report of outstanding work only:

python fccs_report.py

This runs the same document-level check as fccs_verify.py but writes a self-contained, print-friendly HTML file (report_html, default C:\Migration\progress_report.html) that lists only the drawers with missing documents, each with the client name and the specific documents still outstanding. Fully-migrated drawers and drawers with no manifest are omitted (a headline shows how many are done). Open it in any browser and print to PDF to send.

Any outstanding drawer that also appears in crashed.txt (its export was aborted by an FCCS converter crash) is badged CRASHED and sorted to the top, and counted in the header. These are the genuine failures worth spot-checking first — as opposed to benign false positives, where a collapsed "container" document's children exported fine but the container's name lands in the filename's folder field rather than the document field, so it reads as missing.

Config Reference

All scripts read from config.ini (or specify --config path\to\config.ini). config.ini is the git-ignored, per-machine copy auto-created from the tracked config.template.ini; edit config.ini locally.

  • [paths] -- backup_dir, export_dir, output_dir, drawer_id_file, completed_file, ignore_file, crashed_file, log_file, verify_report, report_html, screenshot_dir, manifest_dir, folder_list
  • [timeouts] -- nav_timeout, dialog_timeout, progress_appear, progress_finish, settle, confirm_timeout
  • [controls] -- FCCS window class names and button titles (rarely need changing)