cctbx.xfel GUI/Datasets tab

From cctbx_xfel
Jump to navigation Jump to search

The Datasets tab takes processed runs through to a merged reflection file. A dataset selects runs from a trial by tag and runs a pipeline of tasks over them: optional ensemble refinement, scaling, merging, and optionally Phenix. Each merge creates a new version of the dataset, and new versions keep appearing as further runs finish, so merged data are available while collection is still going on.

Back to cctbx.xfel GUI.

Layout

Each dataset is a box labelled with its id and name, laid out left to right. Inside it:

  • the comment,
  • Tags: the tags that select its runs,
  • Trial: the trial whose results it uses,
  • the numbered pipeline stages, with the resolution limit and reference model name for the scaling and merging stages. A link icon in front of a stage means its task is shared with another dataset; hover for the names.
  • Latest: vNNN: the most recent version merged.
  • The magnifying-glass button opens the full dataset dialog to edit the dataset.
  • The copy button duplicates the dataset: an inactive copy under a new name that shares the same tasks, tags and comment. Use it to merge the same runs with a different resolution limit or reference, after detaching the stages you want to change (see Shared tasks).
  • Active Dataset: while ticked, the job sentinel submits the dataset's jobs. New datasets start inactive.

Below: Filter narrows the list to names containing the text, Show Only Active Datasets hides inactive ones, and New Dataset opens the wizard.

How a dataset runs

With the dataset active and Auto-submit jobs on, the job sentinel does the following on every cycle:

  1. It finds the dataset's runs: the runs in the trial's run blocks that carry the dataset's tags, combined by the tag operator (intersection: all of the tags; union: any of them). A dataset must have at least one tag, and every one of its tags must be in use on the trial's runs, otherwise it is skipped.
  2. For each run it submits the next local task (ensemble refinement, then scaling) as soon as the previous one has finished with status DONE. Indexing is the trial's own integration job, so it must be DONE for the run to enter the pipeline.
  3. Once some runs have finished every local task, it creates a new dataset version, links those runs' jobs to it, and submits the global tasks (merging, then Phenix) for that version. Each time further runs finish, another version is created that includes them. Re-tagging a run out of the dataset also triggers a corrective version.

Versions are numbered from v000 and their output is in <output folder>/<dataset name>/v<NNN>/; the merged file is <name>_v<NNN>_all.mtz. Watch the statistics of each version on the Merging Stats tab.

The new dataset wizard

New Dataset opens a three-step wizard that covers the usual case. Everything it sets can be changed afterwards in the full dialog.

Step 1, Trial & tags
Name (must be unique), the Trial to merge, the Tags selecting its runs and the Tag operator.
Step 2, Structure
Known reference model: give a Reference model file (MTZ or PDB) to scale and post-refine the data against. Unknown structure: give the Unit cell and Space group instead (pre-filled from the trial). With no reference the data are merged as a plain average (a mark1 merge, with no per-image scaling or post-refinement); the resulting MTZ can then serve as the reference model for a second dataset.
Step 3, Options
Include ensemble refinement (on by default). Resolve indexing ambiguity (cosym): ticked automatically when the trial's symmetry has an indexing ambiguity; the label says how many indexing modes exist and how far the cell is from the higher lattice symmetry that causes them (an exact match means the ambiguity is certain, a few tenths of a degree means it is inferred and may or may not be real). High res. limit (d_min, 1.5 Å by default) and Number of bins for the statistics.

Finish creates the dataset, inactive, with an indexing stage, optionally ensemble refinement, scaling and merging.

The dataset dialog

The magnifying-glass button opens the full Dataset Settings dialog.

Identity and run selection

Name, Comment, Trial, Tags and Tag operator as in the wizard. Changing the trial re-reads its unit cell, space group and integration parameters.

Shared parameters (scaling + merging)

These apply to both the scaling and the merging stage so that they agree:

  • Reference: Known reference model with a Reference model path, or No reference model with a Unit cell and Space group.
  • High res. limit (d_min): resolution cut-off for scaling and merging.
  • Resolution scalar (0.96): multiplied by d_min to set the internal limit used when scaling, so that the edge bin is well populated.
  • Number of bins: resolution shells in the statistics tables.
  • Merge anomalous: ticked, Friedel mates are merged. Untick to keep the anomalous signal.

Pipeline

Each stage is a box with an Include this stage tick box (except indexing, which is always there), a link icon (see Shared tasks), and for the PHIL-backed stages an Edit PHIL button that opens the complete parameter set for that program with the simplified controls' values filled in. Anything not covered by the controls, such as the step list, error model or outlier rejection, is set there.

Indexing

Uses the trial's integration results. No settings.

Ensemble refinement

Time-dependent ensemble refinement (cctbx.xfel.time_varying_refinement): the detector geometry and the crystal models of a run are refined together against all of its images, then the images are reintegrated. It improves the agreement between images and so the merging statistics, at the cost of an extra job per run.

  • Pre-split experiments: split the run into chunks locally before submitting, which gives better balanced chunks; otherwise the splitting happens on the cluster, which is faster but counts chunk sizes less accurately.
  • Expand Nave parameters on reintegration: widen the mosaicity model when reintegrating, to recruit more reflections.

The trial's integration parameters (gain in particular) are copied into this stage so that reintegration matches the original integration; they are visible under Edit PHIL.

Scaling

Per-run filtering and scaling against the reference (cctbx.xfel.merge up to the scaling and post-refinement steps):

  • Min. correlation: lattices whose intensities correlate with the reference model below this are rejected (-1 keeps everything).
  • Unit cell rel. length tol.: lattices whose cell edges differ from the reference cell by more than this fraction are rejected (0.1 = 10%).
  • Filter by unit-cell cluster (from Unit Cells tab): instead of the tolerance, accept only lattices within the Mahalanobis cutoff (default 4) of the chosen Cluster component (0 is the largest) of a Cluster file written by the Unit Cell tab. Files in the output folder's cluster/ directory are listed; Browse... finds others.
  • Significance filter sigma: for each image, keep reflections out to the highest resolution bin whose mean I/σ(I) is above this value.

With a reference model this stage scales and post-refines each run; without one it only filters, since the plain average needs neither.

Merging

Merges the scaled runs into one MTZ with statistics. It takes everything from the shared parameters; Edit PHIL exposes the rest (error model, cosym settings, output options). When cosym was requested, this stage runs modify_cosym to resolve the indexing ambiguity before scaling and merging; the cosym space group, number of dimensions and resolution limit are filled in from the trial's symmetry and the shared d_min, and are visible under Edit PHIL.

Phenix

Runs any Phenix command on the merged MTZ. The text box holds the command on the first line and PHIL parameters on the following lines; <PREVIOUS_TASK_MTZ>, <DATASET_NAME> and <DATASET_VERSION> are substituted at run time. Needs a Phenix setup script in Advanced settings. Add it with Add stage at the bottom of the pipeline; a second scaling stage can be added the same way.

Validation

On OK the dialog checks that a reference model or a unit cell and space group are given, that every enabled stage's parameters parse, and that a merge which resolves an indexing ambiguity and then post-refines has a reference model to anchor to (cosym fixes the indexing only up to an overall choice; without a model, remove postrefine from the merging step list instead). Datasets are only written if every check passes.

Shared tasks

A task can belong to more than one dataset. Duplicating a dataset creates such sharing, and the link icon on a stage can link the stage to a task from another dataset, so that two datasets differing only in their runs merge with identical settings. The icon shows a chain when the task is shared; hover to see with whom, click to link to another dataset's task or to unlink (make a private copy).

Changing a shared stage, either with Edit PHIL or through its controls or the shared parameters, asks whether the change should apply to all datasets using the task or detach this dataset onto a private copy of the task first. Cancel leaves everything unchanged. Untick Include this stage on a shared task to drop it from this dataset only.