Files
FCCS/README.md

155 lines
12 KiB
Markdown

# FCCS Data Extraction Tool
Extracts and organizes documents from FileCabinet CS (Thomson Reuters) using GUI automation. Designed to work across multiple client engagements without code changes.
## Prerequisites
- 32-bit Python (must match FCCS architecture)
- `pywinauto` (`pip install pywinauto`)
- `Pillow` (`pip install Pillow`) — required for failure screenshots
- FileCabinet CS installed and restored with client backup data
## Project Layout
| File | Purpose |
|------|---------|
| `config.template.ini` | Tracked template for all paths, timeouts, and FCCS control identifiers |
| `config.ini` | Per-machine working config (git-ignored; auto-created from the template) |
| `fccs_folders.template.txt` | Tracked template for FCCS folder names |
| `fccs_folders.txt` | Per-engagement folder list (git-ignored; auto-created from the template) |
| `fccs_config.py` | Shared config loading, logging, folder pattern building |
| `fccs_scan.py` | Step 1: Scan backup directory for drawer IDs |
| `fccs_export.py` | Step 2: Automate FCCS GUI to export all drawers |
| `fccs_reorganize.py` | Step 3: Parse filenames and rebuild folder structure |
| `fccs_verify.py` | Step 4 (optional): Compare manifests against exported files |
| `fccs_check.py` | Utility: Interactively check if specific drawers exported completely |
| `fccs_report.py` | Utility: Generate a clean client-facing HTML progress report (issues only) |
| `fccs_report_reorganize.py` | Utility: Console report of `_unparsed` leftovers after reorganizing, with suggested missing folder templates |
| `fccs_dump_controls.py` | Utility: Dump control identifiers of an on-screen dialog |
## Setup Per Engagement
1. Edit `config.ini` -- point `backup_dir`, `export_dir`, and `output_dir` at the client's directories. This file is created automatically from `config.template.ini` the first time you run any script (it's git-ignored, so your machine-specific paths never drift the repo). To reset a machine, delete `config.ini` and re-run. Keep shared/default changes in `config.template.ini` (the tracked file).
2. Populate `fccs_folders.txt` -- open FCCS > System Configuration > Document Folders and list each folder template exactly as shown, one per line. Keep the `YYYY` prefix on recurring folders. UltraTax CS folders are detected automatically and do not need to be listed. Like `config.ini`, this file is auto-created from `fccs_folders.template.txt` on first use and is git-ignored — edit it freely per engagement; put only broadly-useful defaults in the tracked template.
## Workflow
### Step 1: Scan Drawers
```
python fccs_scan.py
```
Reads the FCCS data directory (`backup_dir`) and writes all drawer IDs (subfolder names) to `drawer_ids.txt`. It also:
- **Skips non-drawer entries** — only subdirectories are treated as drawers. When `backup_dir` points at FCCS's live data directory (the Restore directory), that folder also contains system folders whose names start with `$` and miscellaneous loose files; both are ignored.
- **Normalizes drawer IDs** — FileCabinet CS ignores `.` characters in drawer IDs, so a folder named `A123.TJ` on disk is searched and displayed in the UI as `A123TJ`. The scan strips dots from folder names when writing `drawer_ids.txt` so the ID matches what FCCS expects (searching the dotted form returns no results). This is done at the source because FCCS embeds the same dot-free ID as the prefix of exported filenames, which the reorganize/verify/report tools all key off. If stripping dots collapses two distinct folders onto one ID, the scan logs a **collision warning** rather than silently dropping a drawer.
- **Flags prefix clashes** — if one drawer ID is a prefix of another (e.g. `02218` and `02218A`), searching the **base** ID in FCCS pops up a selection box that breaks plain automated navigation. These clashes are reported, and the base (shorter) IDs are auto-seeded into `ignore.txt`. The longer, more-specific IDs (`02218A`) search fine and export normally. (Clash detection runs on the dot-normalized IDs, since that's what FCCS actually searches.)
- **Reports ignored drawers** — any IDs listed in `ignore.txt` that exist in this backup are shown as ones the export will skip.
**Ignoring drawers:** The scan creates `ignore.txt` (at `ignore_file`, default `C:\Migration\ignore.txt`) if it doesn't exist and pre-fills it with the clash base IDs — searching those in FCCS shows a selection box that stalls the plain export, so they're skipped by the main export. Open the file and:
- **Delete or comment out** any clash base you'd rather handle fully by hand.
- **Add** any other drawers to skip (e.g. password-protected folders), one ID per line.
Re-running the scan never overwrites your edits — it only appends newly-discovered clashes. `drawer_ids.txt` stays a full inventory; the export skips anything active in the ignore list. Lines starting with `#` are comments.
### Step 2: Export Documents
```
python fccs_export.py
```
Requires FCCS to be open with export destination already configured. Automates the GUI to export every drawer via File > Send To > File. Features:
- **Resumable** -- tracks completed drawers in `completed.txt`; safe to restart
- **Screenshots** -- captures failure states for diagnosis
- **Defensive** -- one bad drawer won't crash the entire run
- **Crash recovery** -- some documents (e.g. UltraTax "Diagnostics" files) crash FCCS's converter (`FileConversionEngine::convert() failed`), which aborts that drawer's export. The script detects the error dialog, screenshots and logs the crashing document, dismisses it, and records the drawer in `crashed.txt` so it's skipped on future runs instead of stalling. Handle crashed drawers manually (export them excluding the poison document); delete a line from `crashed.txt` to retry after fixing.
### Step 3: Reorganize Files
```
python fccs_reorganize.py
```
Parses the flat exported filenames and copies them into an organized structure:
```
output/
ABRAHAM, REBEKAH L./
Tax Documents/
2025/
030126 E-mail re Tax Info.pdf
Billing & Invoices/
2026/
030826 Invoice for 2025 Forms 1040 & IL-1040.pdf
UltraTax CS/
12-31-2008/
2008 Form 1040 Filing Instructions.doc
Permanent File/
Driver's License.pdf
_unparsed/
(files for this client that couldn't be fully parsed, kept for review)
_unparsed/
(only files whose client couldn't be recovered from the filename)
```
Exported filenames follow the format `{drawer_id}_{client_name}_{folder_name}_{creation_date}_{document_name}.ext`. The parser uses folder templates from `fccs_folders.txt` (with `YYYY` expanded via regex) and the creation date (`MM-DD-YYYY`) as anchors to reliably split the underscore-delimited fields. Folder names are decomposed into nested paths that match the FCCS UI structure (e.g. `2025 Tax Documents` becomes `Tax Documents/2025/`). UltraTax CS folders are matched by a built-in pattern.
Files that can't be fully parsed still keep their client: the drawer ID and client name are the first two underscore-delimited tokens and stay recoverable even when the folder/date parse fails, so those files are filed under `{client_name}/_unparsed/` (retaining their original filename). Only files whose client can't be recovered at all fall back to the top-level `_unparsed/`.
Folder names are sanitized for Windows before use — trailing spaces and periods are stripped from each path component, since Windows can't create a directory ending in a space or dot (e.g. a client named `FARR GROUP, P.L.` becomes `FARR GROUP, P.L`). Reorganization is also resilient per file: if one file can't be placed for any reason, the error is logged and counted (reported as `errored` in the summary) and the run continues with the rest rather than aborting.
**Checking for folder template gaps:** After a reorganize run, get a quick internal summary of what didn't parse:
```
python fccs_report_reorganize.py
```
This walks `output_dir`, lists every client that has an `_unparsed` subfolder (plus the top-level `_unparsed`), and — because unparsed files keep their original export filename — recovers the folder field from each name and aggregates them into **suggested template lines** (years generalized to `YYYY`) that can be pasted into `fccs_folders.txt`. Add the missing templates and re-run `fccs_reorganize.py`. Console-only output; unparsed files with no recoverable folder field are counted separately (oddball names, not template gaps).
### Step 4 (optional): Verify Export Completeness
During export, each drawer's document list is captured from the FCCS dialog and saved as a manifest (in `manifest_dir`). These tools compare the manifests against the files actually in the export folder to confirm nothing was missed.
Both compare at the **document level** and share identical matching logic. They account for:
- **Page-splitting** -- a multi-page document exported as `Name Page 1`, `Name Page 2`, … counts as that one document being present.
- **Filename sanitization** -- document titles containing characters illegal in Windows filenames (e.g. `:` `/` `?`) still match the exported files.
Batch-check every drawer that has a manifest:
```
python fccs_verify.py
```
Reports each drawer as `OK` or `INCOMPLETE` (listing the missing documents), plus a summary and any exported drawers that have no manifest. The full report is written to its own file (`verify_report`, default `C:\Migration\verify_report.txt`) as well as the console — separate from the export's `run_log.txt`.
Spot-check specific drawers interactively (e.g. ones the log marked failed, to see whether they actually finished exporting in the background):
```
python fccs_check.py
Drawer ID(s): 08097 18430
```
> Note: because page-splitting means the number of files can't be mapped one-to-one to documents, completeness is judged by document *presence* (is each manifest document represented by at least one exported file), not by exact file counts.
**Client-facing progress report:** For a clean summary to share with the client, generate an HTML report of outstanding work only:
```
python fccs_report.py
```
This runs the same document-level check as `fccs_verify.py` but writes a self-contained, print-friendly HTML file (`report_html`, default `C:\Migration\progress_report.html`) that lists **only** the drawers with missing documents, each with the client name and the specific documents still outstanding. Fully-migrated drawers and drawers with no manifest are omitted (a headline shows how many are done). Open it in any browser and print to PDF to send.
Any outstanding drawer that also appears in `crashed.txt` (its export was aborted by an FCCS converter crash) is badged **CRASHED** and sorted to the top, and counted in the header. These are the genuine failures worth spot-checking first — as opposed to benign false positives, where a collapsed "container" document's children exported fine but the container's name lands in the filename's folder field rather than the document field, so it reads as missing.
## Config Reference
All scripts read from `config.ini` (or specify `--config path\to\config.ini`). `config.ini` is the git-ignored, per-machine copy auto-created from the tracked `config.template.ini`; edit `config.ini` locally.
- **`[paths]`** -- backup_dir, export_dir, output_dir, drawer_id_file, completed_file, ignore_file, crashed_file, log_file, verify_report, report_html, screenshot_dir, manifest_dir, folder_list
- **`[timeouts]`** -- nav_timeout, dialog_timeout, progress_appear, progress_finish, settle, confirm_timeout
- **`[controls]`** -- FCCS window class names and button titles (rarely need changing)