cctbx.xfel GUI

From cctbx_xfel
Revision as of 20:47, 9 October 2026 by Aaron (talk | contribs) (Created page with "{{DISPLAYTITLE:cctbx.xfel GUI}} The '''cctbx.xfel GUI''' is the graphical front end for processing serial crystallography data with cctbx.xfel and DIALS. It watches for new runs as they are collected, submits processing jobs to a cluster or to the local machine, records everything in a database, and shows live feedback on hit rates, indexing rates, unit cells and merging statistics. It can also take a finished set of runs through ensemble refinement, scaling and merging,...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

The cctbx.xfel GUI is the graphical front end for processing serial crystallography data with cctbx.xfel and DIALS. It watches for new runs as they are collected, submits processing jobs to a cluster or to the local machine, records everything in a database, and shows live feedback on hit rates, indexing rates, unit cells and merging statistics. It can also take a finished set of runs through ensemble refinement, scaling and merging, and on to Phenix.

This page explains the ideas behind the GUI and the layout of the main window. The tabs and dialogs are documented on their own pages, and the GUI's Help toolbar button opens the page for whichever tab is showing.

Pages

  • Settings and login: connecting to the database, choosing the facility, the output folder and the queuing system, saving projects.
  • Runs tab: the list of runs and their sample tags; image averaging.
  • Energy tab: FEE spectrometer and ebeam energy calibration (LCLS only).
  • Trials tab: processing parameters (trials) and which runs they apply to (run blocks).
  • Jobs tab: every submitted job and its status; stopping, deleting and restarting jobs.
  • Run Stats tab: live hit rate, indexing rate and resolution plots.
  • Unit Cell tab: unit cell histograms and clustering.
  • Datasets tab: the post-processing pipeline (ensemble refinement, scaling, merging, Phenix).
  • Merging Stats tab: merging statistics and cosym embedding plots per dataset version.
  • Workflows: step-by-step recipes for common tasks.
  • Troubleshooting: what the red lights mean, where the logs are, and common problems.

How the GUI works

Everything lives in a database

The GUI keeps its state in a MySQL database rather than in files. Runs, tags, trials, run blocks, jobs, datasets and the per-image results of processing (spots found, unit cells, resolution, and so on) are all database rows. This has two consequences worth knowing from the start:

  • Several copies of the GUI can be connected to the same experiment at once, for example one submitting jobs at the beamline and one just watching the statistics from an office. See Monitoring mode below.
  • Closing and reopening the GUI loses nothing. Every trial and job is still there when you reconnect.

All tables for one experiment share a prefix called the experiment tag (for example cxi12345_v1). Changing the experiment tag is effectively starting a fresh, empty experiment in the same database. The tag and the connection settings are entered in the Settings dialog when the GUI starts, and cannot be changed without restarting.

The processing jobs themselves write their results to the output folder on disk (see Where the results go below) and log summaries of each image back to the database as they run. The GUI's plots are drawn from those database records.

Runs and tags

A run is the unit of data collection: at LCLS it is an XTC run number; in standalone mode it is a file or folder of images (see Standalone options). The run sentinel polls for new runs and adds them to the database.

Tags are free-form labels attached to runs, typically sample names or conditions (lysozyme, dark, pump_on, ...). A run can carry any number of tags. Tags are what the later stages use to select data: the Run Stats, Unit Cell and Datasets pages all select runs by tag, so tagging runs consistently as they arrive is the single most useful habit when running the GUI live. The Runs tab is where tags are created and applied, and persistent tags can be applied automatically to every new run.

Trials and run blocks

A trial is a set of processing parameters: spot finding, indexing and integration settings for cctbx.xfel.process (or whichever back end is configured). Trials are numbered, and a new trial starts as a copy of the previous one so that tuning parameters is a matter of changing a few values and creating a new trial. Trials are never edited after creation (except for their comment); to change parameters, create a new trial. This keeps the provenance of every result unambiguous: results are stored under the trial number that produced them.

A run block (also called a rungroup, and labelled rg in paths and job listings) is a contiguous range of runs together with the detector and beam settings that apply to them: detector address, beam centre and distance, binning, energy override, reference geometry, pixel mask and so on. A run block can be open (new runs are added to it automatically as they arrive) or have an explicit end run. Run blocks belong to trials, and the same run block can be shared by several trials.

The GUI submits one processing job for every combination of active trial × run block in that trial × run in that block. Marking a trial active is therefore what starts processing, and deactivating it stops new jobs from being submitted (running jobs continue).

Jobs and sentinels

Sentinels are background threads inside the GUI that poll the database and the queuing system. Each has a status light at the bottom of the window (see below). The two most important are toggled from the toolbar:

  • The run sentinel ("Watch for new runs") looks for new runs every 10 seconds, adds them to the database, applies persistent tags, and extends open run blocks to include them.
  • The job sentinel ("Auto-submit jobs") looks every 10 seconds for work that has not been submitted yet: indexing jobs for active trials, and the pipeline stages of active datasets. It submits them using the multiprocessing settings from the Settings dialog (local, Slurm, LSF, ...).

The job monitor queries the queuing system for the status of every unfinished job and records it in the database. It runs while the Jobs or Run Stats tab is showing.

The other sentinels (run stats, unit cell, merging stats, calibration worker) redraw the plots on their tabs and run only while that tab is showing.

Datasets, tasks and versions

Indexing and integration produce one result per image. To get a merged reflection file a dataset is needed. A dataset selects runs (one trial, and a set of tags combined by intersection or union) and runs a pipeline of tasks over them:

  1. Indexing: the trial's own integration results (nothing new is run).
  2. Ensemble refinement (optional): time-dependent ensemble refinement of detector geometry and crystal models, followed by reintegration, per run.
  3. Scaling: filters lattices (unit cell, correlation to a reference) and scales each run against a reference model, or prepares them for a reference-free merge.
  4. Merging: merges every scaled run into one MTZ file, with statistics.
  5. Phenix (optional): any Phenix command run on the merged MTZ.

The first three are local tasks that run once per run and chain automatically: each starts when the previous task for that run has finished. Merging and Phenix are global tasks that run once over all finished runs. Each time merging runs it creates a new dataset version (v000, v001, ...). As more runs finish during an experiment, the job sentinel keeps rolling new versions, so the merged data set grows while data are still being collected. See the Datasets tab.

Projects (saved settings)

All of the settings entered in the Settings dialog (database connection, experiment tag, facility, output folder, queuing) are saved together as a project in ~/.cctbx.xfel/settings_<name>.phil. The GUI starts with the most recently used project, and the Settings dialog can load or save projects by name, so switching between experiments is a matter of loading a different project. The environment variable CCTBX_XFEL_SETTINGS can point at an explicit settings file instead.

The main window

Toolbar

Button What it does
Quit Stops all sentinels, saves the settings and closes the GUI. Jobs already submitted keep running.
Watch for new runs Starts or stops the run sentinel. The icon changes to a pause symbol while it is running and the Run Sentinel light turns green.
Auto-submit jobs Starts or stops the job sentinel. While it runs, every active trial and dataset gets its jobs submitted automatically. The Job Sentinel light turns green.
Settings Opens the Settings dialog. The database connection and experiment tag are locked while the GUI is connected; everything else (output folder, queue, number of processors, facility options) can be changed and takes effect for jobs submitted afterwards.
Large text Toggles larger text and markers in the Run Stats and Unit Cell plots, for reading from across a control room. Takes effect on the next redraw.
Help Opens this documentation in the default web browser, at the page for the tab currently showing. Also available from the Help menu as Online help.

The Watch for new runs, Auto-submit jobs and Settings buttons are hidden in monitoring mode.

Tabs

The tabs, from left to right: Runs, Energy (LCLS only), Trials, Jobs, Run Stats, Unit Cell, Datasets and Merging stats. They follow the order of a typical experiment: tag the runs, set up a trial, watch the jobs and the statistics, build a dataset, and watch the merging statistics.

Switching to a tab starts the sentinel that feeds it and switching away stops it, so only the visible tab costs database queries.

Status lights

A row of lights along the bottom of the window shows the state of each background thread.

Colour Meaning
grey (off) The thread is not running.
green (on) The thread is running and waiting for its next cycle.
yellow The thread is in the middle of a cycle (querying the database or the queue, or redrawing a plot).
red (alert) The thread hit an error and has stopped. The error is printed in the terminal the GUI was started from. See Troubleshooting.
Light Thread Runs when
Run Sentinel Finds new runs Watch for new runs is on
Calib Worker Runs FEE and ebeam calibrations The Energy tab is showing
Job Sentinel Submits jobs Auto-submit jobs is on
Job Monitor Tracks job status in the queue The Jobs or Run Stats tab is showing
Run Stats Sentinel Redraws the Run Stats plot (every 5 s) The Run Stats tab is showing and Auto update is ticked
Unit Cell Sentinel Redraws the unit cell plot (every 15 s) The Unit Cell tab is showing and Auto update is ticked
Merging Stats Sentinel Redraws the merging statistics (every 5 s) The Merging stats tab is showing

Monitoring mode

Starting the GUI with monitoring_mode=True (in the settings file or on the command line) gives a read-only view of an experiment: the Runs, Trials, Jobs and Datasets tabs and the submit buttons are hidden, leaving the Run Stats, Unit Cell and Merging stats tabs. Use it for a second copy of the GUI connected to the same database, so that only one copy submits jobs.

Where the results go

Everything written by jobs goes under the output folder chosen in the Settings dialog.

Path Contents
r0012/003_rg005/ Indexing and integration results for run 12, trial 3, run block 5. Inside: out/ holds the integrated experiments and reflections (*_integrated.expt, *_integrated.refl); stdout/log.out is the job log; all/ holds every image converted to CBF when Dump all images is on (LCLS). In standalone mode with non-numeric run names the folder is named after the run instead of r0012.
r0012/003_rg005/task007/ Results of dataset task 7 (ensemble refinement or scaling) for that run, trial and block.
<dataset name>/v001/ Merging output for version 1 of a dataset: the merged MTZ (<name>_v001_all.mtz), the merging log, and the cosym embedding plots if cosym ran. Phenix tasks write here too.
averages/ Average, maximum and standard deviation images from the Runs tab Average button.
cluster/ Unit cell cluster files written by the Unit Cell tab, for use by the scaling stage of a dataset.
MySql/ Data files of a locally started database server (default location).

The GUI's own files are in ~/.cctbx.xfel/: the project settings bundles and a cfgs/ folder holding the PHIL and locator files generated for every job.

Starting the GUI

cctbx.xfel

opens the login (Settings) dialog followed by the main window. cctbx.xfel -h prints every configuration parameter with its help text; any of them can be given on the command line, for example cctbx.xfel monitoring_mode=True. The GUI prints progress and error messages to the terminal it was started from, so keep that terminal visible.

Older documentation: Brewster et al. (2019), Computational Crystallography Newsletter 10, 22.