l

Deep residual networks for crystallography trained on synthetic data

The use of artificial intelligence to process diffraction images is challenged by the need to assemble large and precisely designed training data sets. To address this, a codebase called Resonet was developed for synthesizing diffraction data and training residual neural networks on these data. Here, two per-pattern capabilities of Resonet are demonstrated: (i) interpretation of crystal resolution and (ii) identification of overlapping lattices. Resonet was tested across a compilation of diffraction images from synchrotron experiments and X-ray free-electron laser experiments. Crucially, these models readily execute on graphics processing units and can thus significantly outperform conventional algorithms. While Resonet is currently utilized to provide real-time feedback for macromolecular crystallography users at the Stanford Synchrotron Radiation Lightsource, its simple Python-based interface makes it easy to embed in other processing frameworks. This work highlights the utility of physics-based simulation for training deep neural networks and lays the groundwork for the development of additional models to enhance diffraction collection and analysis.




l

The High-Pressure Freezing Laboratory for Macromolecular Crystallography (HPMX), an ancillary tool for the macromolecular crystallography beamlines at the ESRF

This article describes the High-Pressure Freezing Laboratory for Macromolecular Crystallography (HPMX) at the ESRF, and highlights new and complementary research opportunities that can be explored using this facility. The laboratory is dedicated to investigating interactions between macromolecules and gases in crystallo, and finds applications in many fields of research, including fundamental biology, biochemistry, and environmental and medical science. At present, the HPMX laboratory offers the use of different high-pressure cells adapted for helium, argon, krypton, xenon, nitrogen, oxygen, carbon dioxide and methane. Important scientific applications of high pressure to macromolecules at the HPMX include noble-gas derivatization of crystals to detect and map the internal architecture of proteins (pockets, tunnels and channels) that allows the storage and diffusion of ligands or substrates/products, the investigation of the catalytic mechanisms of gas-employing enzymes (using oxygen, carbon dioxide or methane as substrates) to possibly decipher intermediates, and studies of the conformational fluctuations or structure modifications that are necessary for proteins to function. Additionally, cryo-cooling protein crystals under high pressure (helium or argon at 2000 bar) enables the addition of cryo-protectant to be avoided and noble gases can be employed to produce derivatives for structure resolution. The high-pressure systems are designed to process crystals along a well defined pathway in the phase diagram (pressure–temperature) of the gas to cryo-cool the samples according to the three-step `soak-and-freeze method'. Firstly, crystals are soaked in a pressurized pure gas atmosphere (at 294 K) to introduce the gas and facilitate its inter­actions within the macromolecules. Samples are then flash-cooled (at 100 K) while still under pressure to cryo-trap macromolecule–gas complexation states or pressure-induced protein modifications. Finally, the samples are recovered after depressurization at cryo-temperatures. The final section of this publication presents a selection of different typical high-pressure experiments carried out at the HPMX, showing that this technique has already answered a wide range of scientific questions. It is shown that the use of different gases and pressure conditions can be used to probe various effects, such as mapping the functional internal architectures of enzymes (tunnels in the haloalkane dehalogenase DhaA) and allosteric sites on membrane-protein surfaces, the interaction of non-inert gases with proteins (oxygen in the hydrogenase ReMBH) and pressure-induced structural changes of proteins (tetramer dissociation in urate oxidase). The technique is versatile and the provision of pressure cells and their application at the HPMX is gradually being extended to address new scientific questions.




l

A web-based dashboard for RELION metadata visualization

Cryo-electron microscopy (cryo-EM) has witnessed radical progress in the past decade, driven by developments in hardware and software. While current software packages include processing pipelines that simplify the image-processing workflow, they do not prioritize the in-depth analysis of crucial metadata, limiting troubleshooting for challenging data sets. The widely used RELION software package lacks a graphical native representation of the underlying metadata. Here, two web-based tools are introduced: relion_live.py, which offers real-time feedback on data collection, aiding swift decision-making during data acquisition, and relion_analyse.py, a graphical interface to represent RELION projects by plotting essential metadata including interactive data filtration and analysis. A useful script for estimating ice thickness and data quality during movie pre-processing is also presented. These tools empower researchers to analyse data efficiently and allow informed decisions during data collection and processing.




l

From femtoseconds to minutes: time-resolved macromolecular crystallography at XFELs and synchrotrons

Over the last decade, the development of time-resolved serial crystallography (TR-SX) at X-ray free-electron lasers (XFELs) and synchrotrons has allowed researchers to study phenomena occurring in proteins on the femtosecond-to-minute timescale, taking advantage of many technical and methodological breakthroughs. Protein crystals of various sizes are presented to the X-ray beam in either a static or a moving medium. Photoactive proteins were naturally the initial systems to be studied in TR-SX experiments using pump–probe schemes, where the pump is a pulse of visible light. Other reaction initiations through small-molecule diffusion are gaining momentum. Here, selected examples of XFEL and synchrotron time-resolved crystallography studies will be used to highlight the specificities of the various instruments and methods with respect to time resolution, and are compared with cryo-trapping studies.




l

Investigation of how gate residues in the main channel affect the catalytic activity of Scytalidium thermophilum catalase

Catalase is an antioxidant enzyme that breaks down hydrogen peroxide (H2O2) into molecular oxygen and water. In all monofunctional catalases the pathway that H2O2 takes to the catalytic centre is via the `main channel'. However, the structure of this channel differs in large-subunit and small-subunit catalases. In large-subunit catalases the channel is 15 Å longer and consists of two distinct parts, including a hydrophobic lower region near the heme and a hydrophilic upper region where multiple H2O2 routes are possible. Conserved glutamic acid and threonine residues are located near the intersection of these two regions. Mutations of these two residues in the Scytalidium thermophilum catalase had no significant effect on catalase activity. However, the secondary phenol oxidase activity was markedly altered, with kcat and kcat/Km values that were significantly increased in the five variants E484A, E484I, T188D, T188I and T188F. These variants also showed a lower affinity for inhibitors of oxidase activity than the wild-type enzyme and a higher affinity for phenolic substrates. Oxidation of heme b to heme d did not occur in most of the studied variants. Structural changes in solvent-chain integrity and channel architecture were also observed. In summary, modification of the main-channel gate glutamic acid and threonine residues has a greater influence on the secondary activity of the catalase enzyme, and the oxidation of heme b to heme d is predominantly inhibited by their conversion to aliphatic and aromatic residues.




l

Structural flexibility of Toscana virus nucleoprotein in the presence of a single-chain camelid antibody

Phenuiviridae nucleoprotein is the main structural and functional component of the viral cycle, protecting the viral RNA and mediating the essential replication/transcription processes. The nucleoprotein (N) binds the RNA using its globular core and polymerizes through the N-terminus, which is presented as a highly flexible arm, as demonstrated in this article. The nucleoprotein exists in an `open' or a `closed' conformation. In the case of the closed conformation the flexible N-terminal arm folds over the RNA-binding cleft, preventing RNA adsorption. In the open conformation the arm is extended in such a way that both RNA adsorption and N polymerization are possible. In this article, single-crystal X-ray diffraction and small-angle X-ray scattering were used to study the N protein of Toscana virus complexed with a single-chain camelid antibody (VHH) and it is shown that in the presence of the antibody the nucleoprotein is unable to achieve a functional assembly to form a ribonucleoprotein complex.




l

AlphaFold-assisted structure determination of a bacterial protein of unknown function using X-ray and electron crystallography

Macromolecular crystallography generally requires the recovery of missing phase information from diffraction data to reconstruct an electron-density map of the crystallized molecule. Most recent structures have been solved using molecular replacement as a phasing method, requiring an a priori structure that is closely related to the target protein to serve as a search model; when no such search model exists, molecular replacement is not possible. New advances in computational machine-learning methods, however, have resulted in major advances in protein structure predictions from sequence information. Methods that generate predicted structural models of sufficient accuracy provide a powerful approach to molecular replacement. Taking advantage of these advances, AlphaFold predictions were applied to enable structure determination of a bacterial protein of unknown function (UniProtKB Q63NT7, NCBI locus BPSS0212) based on diffraction data that had evaded phasing attempts using MIR and anomalous scattering methods. Using both X-ray and micro-electron (microED) diffraction data, it was possible to solve the structure of the main fragment of the protein using a predicted model of that domain as a starting point. The use of predicted structural models importantly expands the promise of electron diffraction, where structure determination relies critically on molecular replacement.




l

Using cryo-EM to understand the assembly pathway of respiratory complex I

Complex I (proton-pumping NADH:ubiquinone oxidoreductase) is the first component of the mitochondrial respiratory chain. In recent years, high-resolution cryo-EM studies of complex I from various species have greatly enhanced the understanding of the structure and function of this important membrane-protein complex. Less well studied is the structural basis of complex I biogenesis. The assembly of this complex of more than 40 subunits, encoded by nuclear or mitochondrial DNA, is an intricate process that requires at least 20 different assembly factors in humans. These are proteins that are transiently associated with building blocks of the complex and are involved in the assembly process, but are not part of mature complex I. Although the assembly pathways have been studied extensively, there is limited information on the structure and molecular function of the assembly factors. Here, the insights that have been gained into the assembly process using cryo-EM are reviewed.




l

A service-based approach to cryoEM facility processing pipelines at eBIC

Electron cryo-microscopy image-processing workflows are typically composed of elements that may, broadly speaking, be categorized as high-throughput workloads which transition to high-performance workloads as preprocessed data are aggregated. The high-throughput elements are of particular importance in the context of live processing, where an optimal response is highly coupled to the temporal profile of the data collection. In other words, each movie should be processed as quickly as possible at the earliest opportunity. The high level of disconnected parallelization in the high-throughput problem directly allows a completely scalable solution across a distributed computer system, with the only technical obstacle being an efficient and reliable implementation. The cloud computing frameworks primarily developed for the deployment of high-availability web applications provide an environment with a number of appealing features for such high-throughput processing tasks. Here, an implementation of an early-stage processing pipeline for electron cryotomography experiments using a service-based architecture deployed on a Kubernetes cluster is discussed in order to demonstrate the benefits of this approach and how it may be extended to scenarios of considerably increased complexity.




l

The crystal structure of mycothiol disulfide reductase (Mtr) provides mechanistic insight into the specific low-molecular-weight thiol reductase activity of Actinobacteria

Low-molecular-weight (LMW) thiols are involved in many processes in all organisms, playing a protective role against reactive species, heavy metals, toxins and antibiotics. Actinobacteria, such as Mycobacterium tuberculosis, use the LMW thiol mycothiol (MSH) to buffer the intracellular redox environment. The NADPH-dependent FAD-containing oxidoreductase mycothiol disulfide reductase (Mtr) is known to reduce oxidized mycothiol disulfide (MSSM) to MSH, which is crucial to maintain the cellular redox balance. In this work, the first crystal structures of Mtr are presented, expanding the structural knowledge and understanding of LMW thiol reductases. The structural analyses and docking calculations provide insight into the nature of Mtrs, with regard to the binding and reduction of the MSSM substrate, in the context of related oxidoreductases. The putative binding site for MSSM suggests a similar binding to that described for the homologous glutathione reductase and its respective substrate glutathione disulfide, but with distinct structural differences shaped to fit the bulkier MSSM substrate, assigning Mtrs as uniquely functioning reductases. As MSH has been acknowledged as an attractive antitubercular target, the structural findings presented in this work may contribute towards future antituberculosis drug development.




l

Characterization of novel mevalonate kinases from the tardigrade Ramazzottius varieornatus and the psychrophilic archaeon Methanococcoides burtonii

Mevalonate kinase is central to the isoprenoid biosynthesis pathway. Here, high-resolution X-ray crystal structures of two mevalonate kinases are presented: a eukaryotic protein from Ramazzottius varieornatus and an archaeal protein from Methanococcoides burtonii. Both enzymes possess the highly conserved motifs of the GHMP enzyme superfamily, with notable differences between the two enzymes in the N-terminal part of the structures. Biochemical characterization of the two enzymes revealed major differences in their sensitivity to geranyl pyrophosphate and farnesyl pyrophosphate, and in their thermal stabilities. This work adds to the understanding of the structural basis of enzyme inhibition and thermostability in mevalonate kinases.




l

Advanced exploitation of unmerged reflection data during processing and refinement with autoPROC and BUSTER

The validation of structural models obtained by macromolecular X-ray crystallography against experimental diffraction data, whether before deposition into the PDB or after, is typically carried out exclusively against the merged data that are eventually archived along with the atomic coordinates. It is shown here that the availability of unmerged reflection data enables valuable additional analyses to be performed that yield improvements in the final models, and tools are presented to implement them, together with examples of the results to which they give access. The first example is the automatic identification and removal of image ranges affected by loss of crystal centering or by excessive decay of the diffraction pattern as a result of radiation damage. The second example is the `reflection-auditing' process, whereby individual merged data items showing especially poor agreement with model predictions during refinement are investigated thanks to the specific metadata (such as image number and detector position) that are available for the corresponding unmerged data, potentially revealing previously undiagnosed instrumental, experimental or processing problems. The third example is the calculation of so-called F(early) − F(late) maps from carefully selected subsets of unmerged amplitude data, which can not only highlight the location and extent of radiation damage but can also provide guidance towards suitable fine-grained parametrizations to model the localized effects of such damage.




l

EMinsight: a tool to capture cryoEM microscope configuration and experimental outcomes for analysis and deposition

The widespread adoption of cryoEM technologies for structural biology has pushed the discipline to new frontiers. A significant worldwide effort has refined the single-particle analysis (SPA) workflow into a reasonably standardized procedure. Significant investments of development time have been made, particularly in sample preparation, microscope data-collection efficiency, pipeline analyses and data archiving. The widespread adoption of specific commercial microscopes, software for controlling them and best practices developed at facilities worldwide has also begun to establish a degree of standardization to data structures coming from the SPA workflow. There is opportunity to capitalize on this moment in the maturation of the field, to capture metadata from SPA experiments and correlate the metadata with experimental outcomes, which is presented here in a set of programs called EMinsight. This tool aims to prototype the framework and types of analyses that could lead to new insights into optimal microscope configurations as well as to define methods for metadata capture to assist with the archiving of cryoEM SPA data. It is also envisaged that this tool will be useful to microscope operators and facilities looking to rapidly generate reports on SPA data-collection and screening sessions.




l

Structural determination and modeling of ciliary microtubules

The axoneme, a microtubule-based array at the center of every cilium, has been the subject of structural investigations for decades, but only recent advances in cryo-EM and cryo-ET have allowed a molecular-level interpretation of the entire complex to be achieved. The unique properties of the nine doublet microtubules and central pair of singlet microtubules that form the axoneme, including the highly decorated tubulin lattice and the docking of massive axonemal complexes, provide opportunities and challenges for sample preparation, 3D reconstruction and atomic modeling. Here, the approaches used for cryo-EM and cryo-ET of axonemes are reviewed, while highlighting the unique opportunities provided by the latest generation of AI-guided tools that are transforming structural biology.




l

Tomo Live: an on-the-fly reconstruction pipeline to judge data quality for cryo-electron tomography workflows

Data acquisition and processing for cryo-electron tomography can be a significant bottleneck for users. To simplify and streamline the cryo-ET workflow, Tomo Live, an on-the-fly solution that automates the alignment and reconstruction of tilt-series data, enabling real-time data-quality assessment, has been developed. Through the integration of Tomo Live into the data-acquisition workflow for cryo-ET, motion correction is performed directly after each of the acquired tilt angles. Immediately after the tilt-series acquisition has completed, an unattended tilt-series alignment and reconstruction into a 3D volume is performed. The results are displayed in real time in a dedicated remote web platform that runs on the microscope hardware. Through this web platform, users can review the acquired data (aligned stack and 3D volume) and several quality metrics that are obtained during the alignment and reconstruction process. These quality metrics can be used for fast feedback for subsequent acquisitions to save time. Parameters such as Alignment Accuracy, Deleted Tilts and Tilt Axis Correction Angle are visualized as graphs and can be used as filters to export only the best tomograms (raw data, reconstruction and intermediate data) for further processing. Here, the Tomo Live algorithms and workflow are described and representative results on several biological samples are presented. The Tomo Live workflow is accessible to both expert and non-expert users, making it a valuable tool for the continued advancement of structural biology, cell biology and histology.




l

Efficient in situ screening of and data collection from microcrystals in crystallization plates

A considerable bottleneck in serial crystallography at XFEL and synchrotron sources is the efficient production of large quantities of homogenous, well diffracting microcrystals. Efficient high-throughput screening of batch-grown microcrystals and the determination of ground-state structures from different conditions is thus of considerable value in the early stages of a project. Here, a highly sample-efficient methodology to measure serial crystallography data from microcrystals by raster scanning within standard in situ 96-well crystallization plates is described. Structures were determined from very small quantities of microcrystal suspension and the results were compared with those from other sample-delivery methods. The analysis of a two-dimensional batch crystallization screen using this method is also described as a useful guide for further optimization and the selection of appropriate conditions for scaling up microcrystallization.




l

Mononuclear binding and catalytic activity of europium(III) and gadolinium(III) at the active site of the model metalloenzyme phosphotriesterase

Lanthanide ions have ideal chemical properties for catalysis, such as hard Lewis acidity, fast ligand-exchange kinetics, high coordination-number preferences and low geometric requirements for coordination. As a result, many small-molecule lanthanide catalysts have been described in the literature. Yet, despite the ability of enzymes to catalyse highly stereoselective reactions under gentle conditions, very few lanthanoenzymes have been investigated. In this work, the mononuclear binding of europium(III) and gadolinium(III) to the active site of a mutant of the model enzyme phosphotriesterase are described using X-ray crystallography at 1.78 and 1.61 Å resolution, respectively. It is also shown that despite coordinating a single non-natural metal cation, the PTE-R18 mutant is still able to maintain esterase activity.




l

Scaling and merging macromolecular diffuse scattering with mdx2

Diffuse scattering is a promising method to gain additional insight into protein dynamics from macromolecular crystallography experiments. Bragg intensities yield the average electron density, while the diffuse scattering can be processed to obtain a three-dimensional reciprocal-space map that is further analyzed to determine correlated motion. To make diffuse scattering techniques more accessible, software for data processing called mdx2 has been created that is both convenient to use and simple to extend and modify. mdx2 is written in Python, and it interfaces with DIALS to implement self-contained data-reduction workflows. Data are stored in NeXus format for software interchange and convenient visualization. mdx2 can be run on the command line or imported as a package, for instance to encapsulate a complete workflow in a Jupyter notebook for reproducible computing and education. Here, mdx2 version 1.0 is described, a new release incorporating state-of-the-art techniques for data reduction. The implementation of a complete multi-crystal scaling and merging workflow is described, and the methods are tested using a high-redundancy data set from cubic insulin. It is shown that redundancy can be leveraged during scaling to correct systematic errors and obtain accurate and reproducible measurements of weak diffuse signals.




l

HEIDI: an experiment-management platform enabling high-throughput fragment and compound screening

The Swiss Light Source facilitates fragment-based drug-discovery campaigns for academic and industrial users through the Fast Fragment and Compound Screening (FFCS) software suite. This framework is further enriched by the option to utilize the Smart Digital User (SDU) software for automated data collection across the PXI, PXII and PXIII beamlines. In this work, the newly developed HEIDI webpage (https://heidi.psi.ch) is introduced: a platform crafted using state-of-the-art software architecture and web technologies for sample management of rotational data experiments. The HEIDI webpage features a data-review tab for enhanced result visualization and provides programmatic access through a representational state transfer application programming interface (REST API). The migration of the local FFCS MongoDB instance to the cloud is highlighted and detailed. This transition ensures secure, encrypted and consistently accessible data through a robust and reliable REST API tailored for the FFCS software suite. Collectively, these advancements not only significantly elevate the user experience, but also pave the way for future expansions and improvements in the capabilities of the system.




l

STOPGAP: an open-source package for template matching, subtomogram alignment and classification

Cryo-electron tomography (cryo-ET) enables molecular-resolution 3D imaging of complex biological specimens such as viral particles, cellular sections and, in some cases, whole cells. This enables the structural characterization of molecules in their near-native environments, without the need for purification or separation, thereby preserving biological information such as conformational states and spatial relationships between different molecular species. Subtomogram averaging is an image-processing workflow that allows users to leverage cryo-ET data to identify and localize target molecules, determine high-resolution structures of repeating molecular species and classify different conformational states. Here, STOPGAP, an open-source package for subtomogram averaging that is designed to provide users with fine control over each of these steps, is described. In providing detailed descriptions of the image-processing algorithms that STOPGAP uses, this manuscript is also intended to serve as a technical resource to users as well as for further community-driven software development.




l

A database overview of metal-coordination distances in metalloproteins

Metalloproteins are ubiquitous in all living organisms and take part in a very wide range of biological processes. For this reason, their experimental characterization is crucial to obtain improved knowledge of their structure and biological functions. The three-dimensional structure represents highly relevant information since it provides insight into the interaction between the metal ion(s) and the protein fold. Such interactions determine the chemical reactivity of the bound metal. The available PDB structures can contain errors due to experimental factors such as poor resolution and radiation damage. A lack of use of distance restraints during the refinement and validation process also impacts the structure quality. Here, the aim was to obtain a thorough overview of the distribution of the distances between metal ions and their donor atoms through the statistical analysis of a data set based on more than 115 000 metal-binding sites in proteins. This analysis not only produced reference data that can be used by experimentalists to support the structure-determination process, for example as refinement restraints, but also resulted in an improved insight into how protein coordination occurs for different metals and the nature of their binding interactions. In particular, the features of carboxylate coordination were inspected, which is the only type of interaction that is commonly present for nearly all metals.




l

Identifying and avoiding radiation damage in macromolecular crystallography

Radiation damage remains one of the major impediments to accurate structure solution in macromolecular crystallography. The artefacts of radiation damage can manifest as structural changes that result in incorrect biological interpretations being drawn from a model, they can reduce the resolution to which data can be collected and they can even prevent structure solution entirely. In this article, we discuss how to identify and mitigate against the effects of radiation damage at each stage in the macromolecular crystal structure-solution pipeline.




l

A small step towards an important goal: fragment screen of the c-di-AMP-synthesizing enzyme CdaA

CdaA is the most widespread diadenylate cyclase in many bacterial species, including several multidrug-resistant human pathogens. The enzymatic product of CdaA, cyclic di-AMP, is a secondary messenger that is essential for the viability of many bacteria. Its absence in humans makes CdaA a very promising and attractive target for the development of new antibiotics. Here, the structural results are presented of a crystallographic fragment screen against CdaA from Listeria monocytogenes, a saprophytic Gram-positive bacterium and an opportunistic food-borne pathogen that can cause listeriosis in humans and animals. Two of the eight fragment molecules reported here were localized in the highly conserved ATP-binding site. These fragments could serve as potential starting points for the development of antibiotics against several CdaA-dependent bacterial species.




l

New insights into the domain of unknown function (DUF) of EccC5, the pivotal ATPase providing the secretion driving force to the ESX-5 secretion system

Type VII secretion (T7S) systems, also referred to as ESAT-6 secretion (ESX) systems, are molecular machines that have gained great attention due to their implications in cell homeostasis and in host–pathogen interactions in mycobacteria. The latter include important human pathogens such as Mycobacterium tuberculosis (Mtb), the etiological cause of human tuberculosis, which constitutes a pandemic accounting for more than one million deaths every year. The ESX-5 system is exclusively found in slow-growing pathogenic mycobacteria, where it mediates the secretion of a large family of virulence factors: the PE and PPE proteins. The secretion driving force is provided by EccC5, a multidomain ATPase that operates using four globular cytosolic domains: an N-terminal domain of unknown function (EccC5DUF) and three FtsK/SpoIIIE ATPase domains. Recent structural and functional studies of ESX-3 and ESX-5 systems have revealed EccCDUF to be an ATPase-like fold domain with potential ATPase activity, the functionality of which is essential for secretion. Here, the crystal structure of the MtbEccC5DUF domain is reported at 2.05 Å resolution, which reveals a nucleotide-free structure with degenerated cis-acting and trans-acting elements involved in ATP binding and hydrolysis. This crystallographic study, together with a biophysical assessment of the interaction of MtbEccC5DUF with ATP/Mg2+, supports the absence of ATPase activity proposed for this domain. It is shown that this degeneration is also present in DUF domains from other ESX and ESX-like systems, which are likely to exhibit poor or null ATPase activity. Moreover, based on an in silico model of the N-terminal region of MtbEccC5DUF, it is hypothesized that MtbEccC5DUF is a degenerated ATPase domain that may have retained the ability to hexamerize. These observations draw attention to DUF domains as structural elements with potential implications in the opening and closure of the membrane pore during the secretion process via their involvement in inter-protomer interactions.




l

What shapes template-matching performance in cryogenic electron tomography in situ?

The detection of specific biological macromolecules in cryogenic electron tomography data is frequently approached by applying cross-correlation-based 3D template matching. To reduce computational cost and noise, high binning is used to aggregate voxels before template matching. This remains a prevalent practice in both practical applications and methods development. Here, the relation between template size, shape and angular sampling is systematically evaluated to identify ribosomes in a ground-truth annotated data set. It is shown that at the commonly used binning, a detailed subtomogram average, a sphere and a heart emoji result in near-identical performance. These findings indicate that with current template-matching practices macromolecules can only be detected with high precision if their shape and size are sufficiently different from the background. Using theoretical considerations, the experimental results are rationalized and it is discussed why primarily low-frequency information remains at high binning and that template matching fails to be accurate because similarly shaped and sized macromolecules have similar low-frequency spectra. These challenges are discussed and potential enhancements for future template-matching methodologies are proposed.




l

High-confidence placement of low-occupancy fragments into electron density using the anomalous signal of sulfur and halogen atoms

Fragment-based drug design using X-ray crystallography is a powerful technique to enable the development of new lead compounds, or probe molecules, against biological targets. This study addresses the need to determine fragment binding orientations for low-occupancy fragments with incomplete electron density, an essential step before further development of the molecule. Halogen atoms play multiple roles in drug discovery due to their unique combination of electronegativity, steric effects and hydrophobic properties. Fragments incorporating halogen atoms serve as promising starting points in hit-to-lead development as they often establish halogen bonds with target proteins, potentially enhancing binding affinity and selectivity, as well as counteracting drug resistance. Here, the aim was to unambiguously identify the binding orientations of fragment hits for SARS-CoV-2 nonstructural protein 1 (nsp1) which contain a combination of sulfur and/or chlorine, bromine and iodine substituents. The binding orientations of carefully selected nsp1 analogue hits were focused on by employing their anomalous scattering combined with Pan-Dataset Density Analysis (PanDDA). Anomalous difference Fourier maps derived from the diffraction data collected at both standard and long-wavelength X-rays were compared. The discrepancies observed in the maps of iodine-containing fragments collected at different energies were attributed to site-specific radiation-damage stemming from the strong X-ray absorption of I atoms, which is likely to cause cleavage of the C—I bond. A reliable and effective data-collection strategy to unambiguously determine the binding orientations of low-occupancy fragments containing sulfur and/or halogen atoms while mitigating radiation damage is presented.




l

Pillar data-acquisition strategies for cryo-electron tomography of beam-sensitive biological samples

For cryo-electron tomography (cryo-ET) of beam-sensitive biological specimens, a planar sample geometry is typically used. As the sample is tilted, the effective thickness of the sample along the direction of the electron beam increases and the signal-to-noise ratio concomitantly decreases, limiting the transfer of information at high tilt angles. In addition, the tilt range where data can be collected is limited by a combination of various sample-environment constraints, including the limited space in the objective lens pole piece and the possible use of fixed conductive braids to cool the specimen. Consequently, most tilt series are limited to a maximum of ±70°, leading to the presence of a missing wedge in Fourier space. The acquisition of cryo-ET data without a missing wedge, for example using a cylindrical sample geometry, is hence attractive for volumetric analysis of low-symmetry structures such as organelles or vesicles, lysis events, pore formation or filaments for which the missing information cannot be compensated by averaging techniques. Irrespective of the geometry, electron-beam damage to the specimen is an issue and the first images acquired will transfer more high-resolution information than those acquired last. There is also an inherent trade-off between higher sampling in Fourier space and avoiding beam damage to the sample. Finally, the necessity of using a sufficient electron fluence to align the tilt images means that this fluence needs to be fractionated across a small number of images; therefore, the order of data acquisition is also a factor to consider. Here, an n-helix tilt scheme is described and simulated which uses overlapping and interleaved tilt series to maximize the use of a pillar geometry, allowing the entire pillar volume to be reconstructed as a single unit. Three related tilt schemes are also evaluated that extend the continuous and classic dose-symmetric tilt schemes for cryo-ET to pillar samples to enable the collection of isotropic information across all spatial frequencies. A fourfold dose-symmetric scheme is proposed which provides a practical compromise between uniform information transfer and complexity of data acquisition.




l

Introduction of the Capsules environment to support further growth of the SBGrid structural biology software collection

The expansive scientific software ecosystem, characterized by millions of titles across various platforms and formats, poses significant challenges in maintaining reproducibility and provenance in scientific research. The diversity of independently developed applications, evolving versions and heterogeneous components highlights the need for rigorous methodologies to navigate these complexities. In response to these challenges, the SBGrid team builds, installs and configures over 530 specialized software applications for use in the on-premises and cloud-based computing environments of SBGrid Consortium members. To address the intricacies of supporting this diverse application collection, the team has developed the Capsule Software Execution Environment, generally referred to as Capsules. Capsules rely on a collection of programmatically generated bash scripts that work together to isolate the runtime environment of one application from all other applications, thereby providing a transparent cross-platform solution without requiring specialized tools or elevated account privileges for researchers. Capsules facilitate modular, secure software distribution while maintaining a centralized, conflict-free environment. The SBGrid platform, which combines Capsules with the SBGrid collection of structural biology applications, aligns with FAIR goals by enhancing the findability, accessibility, interoperability and reusability of scientific software, ensuring seamless functionality across diverse computing environments. Its adaptability enables application beyond structural biology into other scientific fields.




l

Deep-learning map segmentation for protein X-ray crystallographic structure determination

When solving a structure of a protein from single-wavelength anomalous diffraction X-ray data, the initial phases obtained by phasing from an anomalously scattering substructure usually need to be improved by an iterated electron-density modification. In this manuscript, the use of convolutional neural networks (CNNs) for segmentation of the initial experimental phasing electron-density maps is proposed. The results reported demonstrate that a CNN with U-net architecture, trained on several thousands of electron-density maps generated mainly using X-ray data from the Protein Data Bank in a supervised learning, can improve current density-modification methods.




l

Factors affecting macromolecule orientations in thin films formed in cryo-EM

The formation of a vitrified thin film embedded with randomly oriented macromolecules is an essential prerequisite for cryogenic sample electron microscopy. Most commonly, this is achieved using the plunge-freeze method first described nearly 40 years ago. Although this is a robust method, the behaviour of different macromolecules shows great variation upon freezing and often needs to be optimized to obtain an isotropic, high-resolution reconstruction. For a macromolecule in such a film, the probability of encountering the air–water interface in the time between blotting and freezing and adopting preferred orientations is very high. 3D reconstruction using preferentially oriented particles often leads to anisotropic and uninterpretable maps. Currently, there are no general solutions to this prevalent issue, but several approaches largely focusing on sample preparation with the use of additives and novel grid modifications have been attempted. In this study, the effect of physical and chemical factors on the orientations of macromolecules was investigated through an analysis of selected well studied macromolecules, and important parameters that determine the behaviour of proteins on cryo-EM grids were revealed. These insights highlight the nature of the interactions that cause preferred orientations and can be utilized to systematically address orientation bias for any given macromolecule and to provide a framework to design small-molecule additives to enhance sample stability and behaviour.




l

Validation of electron-microscopy maps using solution small-angle X-ray scattering

The determination of the atomic resolution structure of biomacromolecules is essential for understanding details of their function. Traditionally, such a structure determination has been performed with crystallographic or nuclear resonance methods, but during the last decade, cryogenic transmission electron microscopy (cryo-TEM) has become an equally important tool. As the blotting and flash-freezing of the samples can induce conformational changes, external validation tools are required to ensure that the vitrified samples are representative of the solution. Although many validation tools have already been developed, most of them rely on fully resolved atomic models, which prevents early screening of the cryo-TEM maps. Here, a novel and automated method for performing such a validation utilizing small-angle X-ray scattering measurements, publicly available through the new software package AUSAXS, is introduced and implemented. The method has been tested on both simulated and experimental data, where it was shown to work remarkably well as a validation tool. The method provides a dummy atomic model derived from the EM map which best represents the solution structure.




l

A structural role for tryptophan in proteins, and the ubiquitous Trp Cδ1—H⋯O=C (backbone) hydrogen bond

Tryptophan is the most prominent amino acid found in proteins, with multiple functional roles. Its side chain is made up of the hydrophobic indole moiety, with two groups that act as donors in hydrogen bonds: the Nɛ—H group, which is a potent donor in canonical hydrogen bonds, and a polarized Cδ1—H group, which is capable of forming weaker, noncanonical hydrogen bonds. Due to adjacent electron-withdrawing moieties, C—H⋯O hydrogen bonds are ubiquitous in macromolecules, albeit contingent on the polarization of the donor C—H group. Consequently, Cα—H groups (adjacent to the carbonyl and amino groups of flanking peptide bonds), as well as the Cɛ1—H and Cδ2—H groups of histidines (adjacent to imidazole N atoms), are known to serve as donors in hydrogen bonds, for example stabilizing parallel and antiparallel β-sheets. However, the nature and the functional role of interactions involving the Cδ1—H group of the indole ring of tryptophan are not well characterized. Here, data mining of high-resolution (r ≤ 1.5 Å) crystal structures from the Protein Data Bank was performed and ubiquitous close contacts between the Cδ1—H groups of tryptophan and a range of electronegative acceptors were identified, specifically main-chain carbonyl O atoms immediately upstream and downstream in the polypeptide chain. The stereochemical analysis shows that most of the interactions bear all of the hallmarks of proper hydrogen bonds. At the same time, their cohesive nature is confirmed by quantum-chemical calculations, which reveal interaction energies of 1.5–3.0 kcal mol−1, depending on the specific stereochemistry.




l

A snapshot love story: what serial crystallography has done and will do for us

Serial crystallography, born from groundbreaking experiments at the Linac Coherent Light Source in 2009, has evolved into a pivotal technique in structural biology. Initially pioneered at X-ray free-electron laser facilities, it has now expanded to synchrotron-radiation facilities globally, with dedicated experimental stations enhancing its accessibility. This review gives an overview of current developments in serial crystallography, emphasizing recent results in time-resolved crystallography, and discussing challenges and shortcomings.




l

Managing macromolecular crystallographic data with a laboratory information management system

Protein crystallography is an established method to study the atomic structures of macromolecules and their complexes. A prerequisite for successful structure determination is diffraction-quality crystals, which may require extensive optimization of both the protein and the conditions, and hence projects can stretch over an extended period, with multiple users being involved. The workflow from crystallization and crystal treatment to deposition and publication is well defined, and therefore an electronic laboratory information management system (LIMS) is well suited to management of the data. Completion of the project requires key information on all the steps being available and this information should also be made available according to the FAIR principles. As crystallized samples are typically shipped between facilities, a key feature to be captured in the LIMS is the exchange of metadata between the crystallization facility of the home laboratory and, for example, synchrotron facilities. On completion, structures are deposited in the Protein Data Bank (PDB) and the LIMS can include the PDB code in its database, completing the chain of custody from crystallization to structure deposition and publication. A LIMS designed for macromolecular crystallography, IceBear, is available as a standalone installation and as a hosted service, and the implementation of key features for the capture of metadata in IceBear is discussed as an example.




l

The crystal structure of Shethna protein II (FeSII) from Azotobacter vinelandii suggests a domain swap

The Azotobacter vinelandii FeSII protein forms an oxygen-resistant complex with the nitrogenase MoFe and Fe proteins. FeSII is an adrenodoxin-type ferredoxin that forms a dimer in solution. Previously, the crystal structure was solved [Schlesier et al. (2016), J. Am. Chem. Soc. 138, 239–247] with five copies in the asymmetric unit. One copy is a normal adrenodoxin domain that forms a dimer with its crystallographic symmetry mate. The other four copies are in an `open' conformation with a loop flipped out exposing the 2Fe–2S cluster. The open and closed conformations were interpreted as oxidized and reduced, respectively, and the large conformational change in the open configuration allowed binding to nitrogenase. Here, the structure of FeSII was independently solved in the same crystal form. The positioning of the atoms in the unit cell is similar to the earlier report. However, the interpretation of the structure is different. The `open' conformation is interpreted as the product of a crystallization-induced domain swap. The 2Fe–2S cluster is not exposed to solvent, but in the crystal its interacting helix is replaced by the same helix residues from a crystal symmetry mate. The domain swap is complicated, as it is unusual in being in the middle of the protein rather than at a terminus, and it creates arrangements of molecules that can be interpreted in multiple ways. It is also cautioned that crystal structures should be interpreted in terms of the contents of the entire crystal rather than of one asymmetric unit.




l

Protonation of histidine rings using quantum-mechanical methods

Histidine can be protonated on either or both of the two N atoms of the imidazole moiety. Each of the three possible forms occurs as a result of the stereochemical environment of the histidine side chain. In an atomic model, comparing the possible protonation states in situ, looking at possible hydrogen bonding and metal coordination, it is possible to predict which is most likely to be correct. A more direct method is described that uses quantum-mechanical methods to calculate, also in situ, the minimum geometry and energy for comparison, and therefore to more accurately identify the most likely proton­ation state.




l

Crystallographic fragment-binding studies of the Mycobacterium tuberculosis trifunctional enzyme suggest binding pockets for the tails of the acyl-CoA substrates at its active sites and a potential substrate-channeling path between them

The Mycobacterium tuberculosis trifunctional enzyme (MtTFE) is an α2β2 tetrameric enzyme in which the α-chain harbors the 2E-enoyl-CoA hydratase (ECH) and 3S-hydroxyacyl-CoA dehydrogenase (HAD) active sites, and the β-chain provides the 3-ketoacyl-CoA thiolase (KAT) active site. Linear, medium-chain and long-chain 2E-enoyl-CoA molecules are the preferred substrates of MtTFE. Previous crystallographic binding and modeling studies identified binding sites for the acyl-CoA substrates at the three active sites, as well as the NAD binding pocket at the HAD active site. These studies also identified three additional CoA binding sites on the surface of MtTFE that are different from the active sites. It has been proposed that one of these additional sites could be of functional relevance for the substrate channeling (by surface crawling) of reaction intermediates between the three active sites. Here, 226 fragments were screened in a crystallographic fragment-binding study of MtTFE crystals, resulting in the structures of 16 MtTFE–fragment complexes. Analysis of the 121 fragment-binding events shows that the ECH active site is the `binding hotspot' for the tested fragments, with 41 binding events. The mode of binding of the fragments bound at the active sites provides additional insight into how the long-chain acyl moiety of the substrates can be accommodated at their proposed binding pockets. In addition, the 20 fragment-binding events between the active sites identify potential transient binding sites of reaction intermediates relevant to the possible channeling of substrates between these active sites. These results provide a basis for further studies to understand the functional relevance of the latter binding sites and to identify substrates for which channeling is crucial.




l

Cryo2RT: a high-throughput method for room-temperature macromolecular crystallography from cryo-cooled crystals

Advances in structural biology have relied heavily on synchrotron cryo-crystallography and cryogenic electron microscopy to elucidate biological processes and for drug discovery. However, disparities between cryogenic and room-temperature (RT) crystal structures pose challenges. Here, Cryo2RT, a high-throughput RT data-collection method from cryo-cooled crystals that leverages the cryo-crystallography workflow, is introduced. Tested on endothiapepsin crystals with four soaked fragments, thaumatin and SARS-CoV-2 3CLpro, Cryo2RT reveals unique ligand-binding poses, offers a comparable throughput to cryo-crystallography and eases the exploration of structural dynamics at various temperatures.




l

Likelihood-based interactive local docking into cryo-EM maps in ChimeraX

The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.




l

Structural analysis of a ligand-triggered intermolecular disulfide switch in a major latex protein from opium poppy

Several proteins from plant pathogenesis-related family 10 (PR10) are highly abundant in the latex of opium poppy and have recently been shown to play diverse and important roles in the biosynthesis of benzylisoquinoline alkaloids (BIAs). The recent determination of the first crystal structures of PR10-10 showed how large conformational changes in a surface loop and adjacent β-strand are coupled to the binding of BIA compounds to the central hydrophobic binding pocket. A more detailed analysis of these conformational changes is now reported to further clarify how ligand binding is coupled to the formation and cleavage of an intermolecular disulfide bond that is only sterically allowed when the BIA binding pocket is empty. To decouple ligand binding from disulfide-bond formation, each of the two highly conserved cysteine residues (Cys59 and Cys155) in PR10-10 was replaced with serine using site-directed mutagenesis. Crystal structures of the Cys59Ser mutant were determined in the presence of papaverine and in the absence of exogenous BIA compounds. A crystal structure of the Cys155Ser mutant was also determined in the absence of exogenous BIA compounds. All three of these crystal structures reveal conformations similar to that of wild-type PR10-10 with bound BIA compounds. In the absence of exogenous BIA compounds, the Cys59Ser and Cys155Ser mutants appear to bind an unidentified ligand or mixture of ligands that was presumably introduced during expression of the proteins in Escherichia coli. The analysis of conformational changes triggered by the binding of BIA compounds suggests a molecular mechanism coupling ligand binding to the disruption of an intermolecular disulfide bond. This mechanism may be involved in the regulation of biosynthetic reactions in plants and possibly other organisms.




l

Post-translational modifications in the Protein Data Bank

Proteins frequently undergo covalent modification at the post-translational level, which involves the covalent attachment of chemical groups onto amino acids. This can entail the singular or multiple addition of small groups, such as phosphorylation; long-chain modifications, such as glycosylation; small proteins, such as ubiquitination; as well as the interconversion of chemical groups, such as the formation of pyroglutamic acid. These post-translational modifications (PTMs) are essential for the normal functioning of cells, as they can alter the physicochemical properties of amino acids and therefore influence enzymatic activity, protein localization, protein–protein interactions and protein stability. Despite their inherent importance, accurately depicting PTMs in experimental studies of protein structures often poses a challenge. This review highlights the role of PTMs in protein structures, as well as the prevalence of PTMs in the Protein Data Bank, directing the reader to accurately built examples suitable for use as a modelling reference.




l

Surface-mutagenesis strategies to enable structural biology crystallization platforms

A key prerequisite for the successful application of protein crystallography in drug discovery is to establish a robust crystallization system for a new drug-target protein fast enough to deliver crystal structures when the first inhibitors have been identified in the hit-finding campaign or, at the latest, in the subsequent hit-to-lead process. The first crucial step towards generating well folded proteins with a high likelihood of crystallizing is the identification of suitable truncation variants of the target protein. In some cases an optimal length variant alone is not sufficient to support crystallization and additional surface mutations need to be introduced to obtain suitable crystals. In this contribution, four case studies are presented in which rationally designed surface modifications were key to establishing crystallization conditions for the target proteins (the protein kinases Aurora-C, IRAK4 and BUB1, and the KRAS–SOS1 complex). The design process which led to well diffracting crystals is described and the crystal packing is analysed to understand retrospectively how the specific surface mutations promoted successful crystallization. The presented design approaches are routinely used in our team to support the establishment of robust crystallization systems which enable structure-guided inhibitor optimization for hit-to-lead and lead-optimization projects in pharmaceutical research.




l

Microcrystal electron diffraction structure of Toll-like receptor 2 TIR-domain-nucleated MyD88 TIR-domain higher-order assembly

Eukaryotic TIR (Toll/interleukin-1 receptor protein) domains signal via TIR–TIR interactions, either by self-association or by interaction with other TIR domains. In mammals, TIR domains are found in Toll-like receptors (TLRs) and cytoplasmic adaptor proteins involved in pro-inflammatory signaling. Previous work revealed that the MAL TIR domain (MALTIR) nucleates the assembly of MyD88TIR into crystalline arrays in vitro. A microcrystal electron diffraction (MicroED) structure of the MyD88TIR assembly has previously been solved, revealing a two-stranded higher-order assembly of TIR domains. In this work, it is demonstrated that the TIR domain of TLR2, which is reported to signal as a heterodimer with either TLR1 or TLR6, induces the formation of crystalline higher-order assemblies of MyD88TIR in vitro, whereas TLR1TIR and TLR6TIR do not. Using an improved data-collection protocol, the MicroED structure of TLR2TIR-induced MyD88TIR microcrystals was determined at a higher resolution (2.85 Å) and with higher completeness (89%) compared with the previous structure of the MALTIR-induced MyD88TIR assembly. Both assemblies exhibit conformational differences in several areas that are important for signaling (for example the BB loop and CD loop) compared with their monomeric structures. These data suggest that TLR2TIR and MALTIR interact with MyD88 in an analogous manner during signaling, nucleating MyD88TIR assemblies uni­directionally.




l

Comparison of two crystal polymorphs of NowGFP reveals a new conformational state trapped by crystal packing

Crystal polymorphism serves as a strategy to study the conformational flexibility of proteins. However, the relationship between protein crystal packing and protein conformation often remains elusive. In this study, two distinct crystal forms of a green fluorescent protein variant, NowGFP, are compared: a previously identified monoclinic form (space group C2) and a newly discovered ortho­rhombic form (space group P212121). Comparative analysis reveals that both crystal forms exhibit nearly identical linear assemblies of NowGFP molecules interconnected through similar crystal contacts. However, a notable difference lies in the stacking of these assemblies: parallel in the monoclinic form and perpendicular in the orthorhombic form. This distinct mode of stacking leads to different crystal contacts and induces structural alteration in one of the two molecules within the asymmetric unit of the orthorhombic crystal form. This new conformational state captured by orthorhombic crystal packing exhibits two unique features: a conformational shift of the β-barrel scaffold and a restriction of pH-dependent shifts of the key residue Lys61, which is crucial for the pH-dependent spectral shift of this protein. These findings demonstrate a clear connection between crystal packing and alternative conformational states of proteins, providing insights into how structural variations influence the function of fluorescent proteins.




l

Robust and automatic beamstop shadow outlier rejection: combining crystallographic statistics with modern clustering under a semi-supervised learning strategy

During the automatic processing of crystallographic diffraction experiments, beamstop shadows are often unaccounted for or only partially masked. As a result of this, outlier reflection intensities are integrated, which is a known issue. Traditional statistical diagnostics have only limited effectiveness in identifying these outliers, here termed Not-Excluded-unMasked-Outliers (NEMOs). The diagnostic tool AUSPEX allows visual inspection of NEMOs, where they form a typical pattern: clusters at the low-resolution end of the AUSPEX plots of intensities or amplitudes versus resolution. To automate NEMO detection, a new algorithm was developed by combining data statistics with a density-based clustering method. This approach demonstrates a promising performance in detecting NEMOs in merged data sets without disrupting existing data-reduction pipelines. Re-refinement results indicate that excluding the identified NEMOs can effectively enhance the quality of subsequent structure-determination steps. This method offers a prospective automated means to assess the efficacy of a beamstop mask, as well as highlighting the potential of modern pattern-recognition techniques for automating outlier exclusion during data processing, facilitating future adaptation to evolving experimental strategies.




l

Utilizing anomalous signals for element identification in macromolecular crystallography

AlphaFold2 has revolutionized structural biology by offering unparalleled accuracy in predicting protein structures. Traditional methods for determining protein structures, such as X-ray crystallography and cryo-electron microscopy, are often time-consuming and resource-intensive. AlphaFold2 provides models that are valuable for molecular replacement, aiding in model building and docking into electron density or potential maps. However, despite its capabilities, models from AlphaFold2 do not consistently match the accuracy of experimentally determined structures, need to be validated experimentally and currently miss some crucial information, such as post-translational modifications, ligands and bound ions. In this paper, the advantages are explored of collecting X-ray anomalous data to identify chemical elements, such as metal ions, which are key to understanding certain structures and functions of proteins. This is achieved through methods such as calculating anomalous difference Fourier maps or refining the imaginary component of the anomalous scattering factor f''. Anomalous data can serve as a valuable complement to the information provided by AlphaFold2 models and this is particularly significant in elucidating the roles of metal ions.




l

Structural studies of β-glucosidase from the thermophilic bacterium Caldicellulosiruptor saccharolyticus

β-Glucosidase from the thermophilic bacterium Caldicellulosiruptor saccharo­lyticus (Bgl1) has been denoted as having an attractive catalytic profile for various industrial applications. Bgl1 catalyses the final step of in the decomposition of cellulose, an unbranched glucose polymer that has attracted the attention of researchers in recent years as it is the most abundant renewable source of reduced carbon in the biosphere. With the aim of enhancing the thermostability of Bgl1 for a broad spectrum of biotechnological processes, it has been subjected to structural studies. Crystal structures of Bgl1 and its complex with glucose were determined at 1.47 and 1.95 Å resolution, respectively. Bgl1 is a member of glycosyl hydrolase family 1 (GH1 superfamily, EC 3.2.1.21) and the results showed that the 3D structure of Bgl1 follows the overall architecture of the GH1 family, with a classical (β/α)8 TIM-barrel fold. Comparisons of Bgl1 with sequence or structural homologues of β-glucosidase reveal quite similar structures but also unique structural features in Bgl1 with plausible functional roles.




l

CHiMP: deep-learning tools trained on protein crystallization micrographs to enable automation of experiments

A group of three deep-learning tools, referred to collectively as CHiMP (Crystal Hits in My Plate), were created for analysis of micrographs of protein crystallization experiments at the Diamond Light Source (DLS) synchrotron, UK. The first tool, a classification network, assigns images into categories relating to experimental outcomes. The other two tools are networks that perform both object detection and instance segmentation, resulting in masks of individual crystals in the first case and masks of crystallization droplets in addition to crystals in the second case, allowing the positions and sizes of these entities to be recorded. The creation of these tools used transfer learning, where weights from a pre-trained deep-learning network were used as a starting point and repurposed by further training on a relatively small set of data. Two of the tools are now integrated at the VMXi macromolecular crystallography beamline at DLS, where they have the potential to absolve the need for any user input, both for monitoring crystallization experiments and for triggering in situ data collections. The third is being integrated into the XChem fragment-based drug-discovery screening platform, also at DLS, to allow the automatic targeting of acoustic compound dispensing into crystallization droplets.




l

The success rate of processed predicted models in molecular replacement: implications for experimental phasing in the AlphaFold era

The availability of highly accurate protein structure predictions from AlphaFold2 (AF2) and similar tools has hugely expanded the applicability of molecular replacement (MR) for crystal structure solution. Many structures can be solved routinely using raw models, structures processed to remove unreliable parts or models split into distinct structural units. There is therefore an open question around how many and which cases still require experimental phasing methods such as single-wavelength anomalous diffraction (SAD). Here, this question is addressed using a large set of PDB depositions that were solved by SAD. A large majority (87%) could be solved using unedited or minimally edited AF2 predictions. A further 18 (4%) yield straightforwardly to MR after splitting of the AF2 prediction using Slice'N'Dice, although different splitting methods succeeded on slightly different sets of cases. It is also found that further unique targets can be solved by alternative modelling approaches such as ESMFold (four cases), alternative MR approaches such as ARCIMBOLDO and AMPLE (two cases each), and multimeric model building with AlphaFold-Multimer or UniFold (three cases). Ultimately, only 12 cases, or 3% of the SAD-phased set, did not yield to any form of MR tested here, offering valuable hints as to the number and the characteristics of cases where experimental phasing remains essential for macromolecular structure solution.




l

EMhub: a web platform for data management and on-the-fly processing in scientific facilities

Most scientific facilities produce large amounts of heterogeneous data at a rapid pace. Managing users, instruments, reports and invoices presents additional challenges. To address these challenges, EMhub, a web platform designed to support the daily operations and record-keeping of a scientific facility, has been introduced. EMhub enables the easy management of user information, instruments, bookings and projects. The application was initially developed to meet the needs of a cryoEM facility, but its functionality and adaptability have proven to be broad enough to be extended to other data-generating centers. The expansion of EMHub is enabled by the modular nature of its core functionalities. The application allows external processes to be connected via a REST API, automating tasks such as folder creation, user and password generation, and the execution of real-time data-processing pipelines. EMhub has been used for several years at the Swedish National CryoEM Facility and has been installed in the CryoEM center at the Structural Biology Department at St. Jude Children's Research Hospital. A fully automated single-particle pipeline has been implemented for on-the-fly data processing and analysis. At St. Jude, the X-Ray Crystallography Center and the Single-Molecule Imaging Center have already expanded the platform to support their operational and data-management workflows.