Built with Claude: Life Sciences · Research Track

The cis-regulatory code of valve endothelial cells

Tanvi Sinha & Jan Lebert
UCSF
Cardiovascular Research Institute
University of California, San Francisco

Heart-valve disease affects roughly one in forty people and is a leading reason for cardiac surgery in adults, yet surgical repair remains the primary treatment. Bicuspid aortic valve, the most common congenital heart defect and present in roughly 1% of the population, predisposes to valve calcification and failure decades later, and many adult valve pathologies can be traced to the developmental program that first established the tissue. Understanding valve disease therefore requires understanding valve development: the transcription factors that construct the valve are, when perturbed, the same factors that drive its degeneration, and could in turn serve as points of therapeutic intervention.

A major barrier to understanding both valve development and disease is our incomplete knowledge of the gene-regulatory program that builds and maintains the valve. This project set out to uncover that program at the level of its regulatory DNA: to define the cis-regulatory code of valve endothelial cells.

01 · BackgroundValve development is a route into valve disease

Heart valves are lined by specialized endothelial cells, which have distinct gene-expression profiles and a heightened ability to transduce the mechanical force of flowing blood into the transcriptional program that shapes and maintains the leaflets. Disruption of this program during development produces congenital valve malformation, while its progressive dysregulation with age underlies calcific and degenerative valve disease.

Although substantial progress has been made in identifying the signaling pathways involved in valve development and disease, the upstream transcriptional circuitry and regulatory elements that control valve-specific gene expression remain largely unknown.

Research question
What is the cis-regulatory code of valve endothelial cells — the transcription-factor motifs and regulatory elements that specify valve-endothelial identity — and which of its elements are functionally required?
We addressed this by integrating developmental and adult single-cell multiomic data across two species, and by training a sequence model to identify the most important motifs and predict the effect of mutating them.
On the role of Claude Science

Tanvi Sinha is an experimental cardiovascular biologist. The computational work behind this project — from single-cell genomics to training deep-learning models — is what she would normally collaborate on with a computational specialist like her teammate Jan Lebert. In this project we test whether Claude Science can be that specialist. Tanvi directed the work throughout; Claude Science proposed the analyses, wrote and ran the code, and interpreted the output together with Tanvi and Jan. We also used it to try approaches a wet-lab group would rarely build alone — interactive dashboards, and deep-learning models trained on a GPU. Every figure and result here was produced using Claude Science, from public datasets it located itself.

02 · ApproachDevelopmental and adult multiomics across two species

We characterised valve-endothelial gene regulation by combining several orthogonal lines of evidence rather than testing candidate factors individually: transcription-factor expression in valve endothelial cells (single-cell RNA-seq from mouse and zebrafish hearts); chromatin accessibility defining the active regulatory elements in those cells (single-cell ATAC-seq, spanning both embryonic development and the adult heart); and the sequence grammar extracted by a deep-learning model trained to predict that accessibility from DNA alone. Sampling both developmental and adult stages across two vertebrate species isolates the conserved core of the regulatory code from stage- and species-specific components.

03 · Cell atlasA conserved valve-endothelial program in mouse and zebrafish

Valve endothelial cells were resolved from single-cell RNA-seq in both mice and zebrafish and profiled for transcription-factor expression. Eleven factor families are enriched in valve endothelium in both mouse and zebrafish, indicating a regulatory program conserved across ~400 million years of vertebrate divergence. Among these, klf2a is expressed in 94% of zebrafish valve endothelial cells, consistent with its established role as a flow-responsive valve-endothelial regulator.

Cardiac cell atlas (mouse). Valve endothelial cells resolved from embryonic-heart single-cell RNA-seq across three developmental stages (E8.5, E10.5 and E16.5), marked by Prox1, Foxc1, Gata5, Nfatc1, Hand2 and Klf2.

04 · Regulatory elementsA catalog of valve-specific cis-regulatory elements

Single-cell chromatin-accessibility data were used to identify the cis-regulatory elements (CREs) that distinguish valve endothelium from all other endothelial cells, across both embryonic development and the adult heart. In the developing mouse valve this yielded 1,042 valve-specific CREs, and the program becomes more sharply defined as the leaflets remodel: roughly nine times as many valve-specific elements are present at the later stage as at the earlier cushion stage. Motif enrichment within these elements identified transcription-factor families including NFAT and KLF, known regulators of valve development, as well as cardiac transcription factors such as GATA and TBX.

Valve-specific CREs. Valve endocardium (red) separated from other endothelium and cushion mesenchyme; 940 valve-specific elements at the leaflet-remodeling stage, 102 at the earlier stage.
Motif enrichment within the valve CREs. NFAT is most enriched at both stages, with FOX, SOX and GATA early and KLF, NFI and AP-1 added as the valve matures.

05 · Sequence modelA deep-learning model identifies the important motifs and their mutational effects

To move from correlation to a model of the regulatory sequence itself, we trained a ChromBPNet model to predict valve chromatin accessibility directly from DNA; it reaches r = 0.74 on held-out chromosomes. Interpreting the trained model recovers the governing transcription-factor motifs de novo, without any motif supplied in advance: the endothelial ETS factors, the flow-responsive KLF family, NFI and AP-1. Because the model reads accessibility from sequence, it also predicts the effect of mutating any single base (in-silico mutagenesis), ranking which positions within each element are functionally important. This yields, for each CRE, a prioritised set of bases and motifs whose disruption is predicted to change accessibility. These are directly testable in vivo, allowing the predicted mutational sensitivities to be compared against experimental perturbation.

Motifs learned from sequence alone. The transcription-factor motifs the model recovered de novo, including ETS, KLF, NFI and AP-1/Jun.
In-silico mutagenesis of a mouse valve CRE. Per-base predicted effect of every substitution across an example element (mouse, chr4); the model localises the functionally important positions to specific motifs (NFY, ETS, KLF). In the lower saturation-mutagenesis heatmap, each column is a position and each row an alternative allele: red marks substitutions predicted to increase accessibility and blue those predicted to decrease it, giving a mutational map that can be tested experimentally.

Three independent lines of evidence converge on why these elements are valve-specific rather than generically open. First, the sequence model: in-silico mutagenesis concentrates its highest per-base importance on ETS motifs, with KLF and NFY as secondary contributors, so the model reads valve accessibility primarily off an ETS grammar it was never given. Second, the real mouse single-cell data: scanning the valve CREs against a dinucleotide-shuffled null, an ETS motif is present in 82% of them (versus 42% expected), and more than half carry two or more ETS sites in tandem (53% versus 10%; ~1.8 sites per CRE), with this density highest in the earliest valve cells (91% of E10.5 CREs). Specificity is then built combinatorially, ETS co-occurring with the flow-responsive KLF family (~10-fold over null), with AP-1 (~8-fold) and with GATA and FOX. Third, the same analysis in our adult zebrafish scATAC recovers the conserved core: among ETS-containing valve CREs, KLF is the significantly enriched partner (odds ratio 1.8, p = 0.04), pinpointing a mechanosensitive ETS/KLF module as the cross-species heart of the code, while the more elaborate AP-1 and NFAT partnerships appear to be features of the developing mouse valve. Together, a model trained only on sequence and two real multiomic datasets in two species independently point to clustered ETS sites read out by KLF as the molecular basis of valve-endothelial enhancer specificity.

06 · Cross-species conservationOverlap between the mouse and zebrafish regulatory code

We also applied the same pipeline to zebrafish valve endothelium to allow a direct cross-species comparison. The mouse and fish valve-CRE sets share transcription-factor motif usage well above chance: an ETS + AP-1 + KLF core recurs in both, centred on the klf2a-associated mechanosensitive program, and the valve-CRE sets from two independent zebrafish datasets overlap at 3.8× the frequency expected by chance. The concordance of an unbiased mouse analysis and an independent zebrafish analysis indicates that the identified code reflects conserved valve-endothelial regulation between species.

07 · Experimental validationRecovered CREs include two enhancers already validated in vivo

As an independent test of the pipeline, we examined whether its unbiased, genome-wide output recovered regulatory elements of known function. The valve-specific CREs identified in the adult zebrafish valve endocardium include two enhancers that our work has independently tested in vivo, one near the valve-enriched transcription factor klf2a (~40 kb upstream) and one near mir126a (~1.5 kb upstream) (data shown below), both confirmed experimentally to be valve-specific. Their recovery by the analysis, without any prior knowledge of them, indicates that the computational pipeline built with Claude Science can identify bona fide valve regulatory elements, and supports the validity of the remaining, not-yet-tested CREs as candidates for in vivo testing. It also supports using this compendium of identified CREs to robustly map the transcriptional circuitry that controls their activity.

mir126a locus, two independent assays on one axis (danRer11). Top panel: FANS bulk chromatin accessibility in the developing zebrafish heart endothelium (fli1a+ endothelial over non-endothelial); bottom panel: zebrafish adult-heart single-cell ATAC (valve-EC over other endothelium). The ~1.5 kb upstream valve enhancer is more accessible in the endothelial and valve-EC profiles.

We further asked whether the sequence model captures the regulatory logic of these two enhancers. Because the model was trained only on mouse chromatin, applying it to the zebrafish enhancer sequences is a cross-species test: it is informative for conserved motifs but not for absolute, species-specific effect sizes. In-silico mutagenesis of both enhancers with the mouse-trained model concentrates its highest base-importance scores on ETS motifs, mapped independently by FIMO against the JASPAR database, and this is the same endothelial ETS grammar the model recovered de novo in mouse. That a mouse-trained model localises importance to the conserved ETS sites of a zebrafish enhancer it has never seen is consistent with the deep conservation of the valve-endothelial code, and provides an orthogonal, sequence-level line of support for these elements. The paired ETS motifs the model identifies in the klf2a enhancer are themselves highly conserved across six other vertebrate species, extending back to the ancestral coelacanth, underscoring the importance of this transcription-factor motif.

Cross-species in-silico mutagenesis of the two validated enhancers. Per-base predicted effect of every substitution across each zebrafish enhancer, scored with the mouse-trained model (colour = reference base). Shaded spans are ETS motifs called independently by FIMO (JASPAR2022, q < 0.05); the model's highest-importance positions coincide with them in both enhancers. Interpretable for conserved motifs only; absolute effect magnitudes are not comparable across species.

08 · CollaborationClaude Science as a computational collaborator

Working with Claude Science as that specialist was, overall, excellent: Tanvi rarely ran into problems and needed little help from Jan. We especially liked its Plans, which scoped each multi-step analysis before it ran, so we could see and adjust the approach up front. We also liked that it would naturally offer to build reusable capability as it worked — Skills that capture an entire workflow, such as chrombpnet-scatac (the whole scATAC-to-ChromBPNet pipeline), alongside its built-in skills for ArchR and scanpy, remote GPU compute over SSH, and genome liftover; and custom specialist agents such as REGULOME, which we configured for single-cell multiomics and reused throughout.

The pipeline captured as a reusable skill. The chrombpnet-scatac skill published to the project's catalog within the Claude Science workbench, so the scATAC-to-model workflow can be re-run in one step.

Tanvi very much liked its interactive-dashboard creation skills. Rather than leaving each analysis as a static set of figures, Claude Science could put a point-and-click graphical interface on top of a project — a way to explore the data and the model without writing code. Over the project it built several: a live in-silico mutagenesis browser that runs the ChromBPNet model on the GPU and scores every single-base substitution in a chosen valve-EC element in real time; a CRE explorer linking model-predicted activity against measured specificity for all 1,042 elements, filterable by stage, region and motif family; a TF motif atlas of the motifs the model discovered de novo, as sequence logos; a UMAP cell explorer for recolouring the cardiac atlas by cell type, stage or marker gene; and a cross-species and human valve-disease GWAS dashboard.

Interactive tools built by Claude Science. A hub of the dashboards created over the project — the live GPU-backed in-silico mutagenesis browser, the CRE explorer, the TF motif atlas, the UMAP cell explorer, and the cross-species / GWAS dashboard — each a point-and-click interface a wet-lab scientist can drive directly, no code required.
The live in-silico mutagenesis browser. Pick a stage and a valve-EC element, and the model scores every single-base substitution in real time, with per-base importance and the full mutagenesis heatmap.

One capability stood out as unusually practical: requesting changes to figures made earlier, even ones produced in a different thread. We could point to a plot from days before, ask for a relabelled axis or a new annotation, and Claude Science would work out where all the underlying code and data came from, recreate the figure with the changes, and even find and update the related figures built from the same analysis — with no need for us to track down the original scripts or intermediate files.

A few things would make it a still better collaborator for a wet-lab group. Its outputs are called artifacts, but in biology an “artifact” usually means an error or spurious feature in the data or a figure, so a term such as attachments would be clearer for newcomers. Working as a team is also still hard: we would like to share a project or report directly, so that one of us can pick up where another left off. And we could not make Claude Science run or serve from a remote machine (--host 0.0.0.0 did not take).

A few notes on the tooling itself. Claude Science runs automated reviewer checkpoints that sanity-check results as they are produced and caught small mistakes for us; we would like them to go beyond nitpicks and also weigh in on the bigger picture and the plan. It also, fairly often, produces figures with overlapping or blocked text and annotations, and should look more closely at the rendered figures to catch and fix them. And at the model level, we could not select Fable 5 in Claude Science, even for non-biological work — and its safeguards would trigger on very benign biological questions and everyday research, which was a recurring frustration.

09 · LimitationsWhere the analysis — and Claude Science — needed support

The workflow was productive but not autonomous; several points required scientist oversight or remain limited:

10 · Next stepsWhere we take this next

The immediate priority is experimental. First, we will test several of the novel valve-endothelial-specific enhancers identified from the single-cell ATAC-seq analyses in mouse and zebrafish, to validate whether the computational results match enhancer activity in vivo. For every valve-endothelial CRE that is active in vivo, we will then run in-silico mutagenesis: the model gives a ranked set of bases and motifs whose disruption it predicts will change accessibility, and these are directly testable. We will take the highest-confidence predictions first — the ETS, KLF and AP-1 sites in the validated klf2a and mir126a enhancers — into reporter and CRISPR mutagenesis in zebrafish and mouse, comparing the measured effect of each substitution against the model's prediction.

We would also like to test whether these enhancers are functionally required in vivo for valve development and maintenance. In parallel, it would be interesting to test whether another sequence model, such as Akita, can predict changes in 3D conformation that disrupt enhancer–promoter interactions and affect gene expression, especially for distally located enhancers. This would let us identify enhancer interactions with genes not previously linked to valve endothelial cells, and in turn uncover new players in valve biology.

Finally, we plan to confirm the predicted transcription-factor binding with direct assays (ChIP-seq, footprinting) rather than motif inference alone; build a conserved, multi-species CRE set large enough to give human valve-disease genetics real power, addressing the underpowered GWAS reported above; and deepen the sequence model with more valve-EC cells and additional stages and species.

11 · Plot galleryMore plots we liked

Beyond the figures above, here are some other plots Claude Science produced for us over the course of the project. Click any thumbnail for the full-resolution image.

A note on method. Transcription-factor motif matches are binding predictions, not measured occupancy. The ChromBPNet model predicts chromatin accessibility from DNA sequence and makes no clinical claim. The human valve-disease GWAS test returned a transparent underpowered null. Datasets are public (GEO/ENCODE); analysis was performed during the project. Open-source.

References. Valve-disease prevalence (~2.5%, ≈ 1 in 40): Nkomo VT, Gardin JM, Skelton TN, Gottdiener JS, Scott CG, Enriquez-Sarano M. Burden of valvular heart diseases: a population-based study. Lancet 2006;368(9540):1005–1011 (PMID 16980116). Bicuspid aortic valve prevalence (~1%): Siu SC, Silversides CK. Bicuspid aortic valve disease. J Am Coll Cardiol 2010;55(25):2789–2800.