Repository navigation
Conversation
…on the inferred reference sequence Previous builds inferred the tree's root sequence independently of the dataset reference, causing mismatched mutations at the root and a Nextclade preprocessing error. The tree is now rebuilt with ancestral reconstruction anchored to the reference (inferred ancestral) sequence.
- change order of references in Nextclade results table - specify correct order in README - remove non-VP1 mutations for VP1-only clades (previous issue fixed)
|
Note ✨ AI is responding on behalf of Ivan First review of PR #466, adding the Verdict: no blocking issues. The dataset is structurally sound and scientifically well founded: coordinates, CDS divisibility, GFF3/tree agreement, generated output, registration, and a live Nextclade CLI run all pass. There are 9 non-blocking items (5 medium, 4 low), all concerning documentation accuracy, user-facing links, and reference-tree presentation rather than analysis correctness. TestingObservedWhat the dataset is, and where it diverges from its description [click to expand]The PR adds one whole-genome Coxsackievirus A10 dataset at The PR description contradicts the shipped design: it says the dataset is "rooted on the Kowalik reference sequence" and adds "the Kowalik reference genome sequence in BackgroundCoxsackievirus A10: taxonomy, genome, and the A-H genotype scheme [click to expand]Taxonomy and reference strainCoxsackievirus A10 is a serotype within the species Enterovirus alphacoxsackie (the ICTV binomial adopted in the 2023 Master Species List; formerly Enterovirus A), genus Enterovirus, family Picornaviridae, and is a positive-sense single-stranded RNA virus, Baltimore group IV (NCBI Taxonomy 42769). The reference is the Kowalik prototype, AY421767.1, 7409 nt, sequenced in the complete-genome survey of the species by Oberste et al., J Gen Virol 2004. There is no RefSeq genome for CVA10. Genome organizationThe genome is a single ~7.4 kb ORF: Clade and lineage classificationJi et al., Sci Rep 2018 defines genotypes A-G on the 894 nt VP1 gene with a 14.97% between-genotype divergence threshold, splits C into C1/C2, and reports H and I as candidates detectable only on partial VP1. Yang et al., J Med Virol 2025 works with the eight-genotype A-H scheme this dataset implements. Geographically, A is the prototype alone, B was a ceased Chinese lineage, C (especially C2) dominates Asia, D is the European lineage (also entering China after 2016), E is West/Central African, F is Indian, and G was first described from Taiwan. Clinically CVA10 is an established cause of hand-foot-and-mouth disease, herpangina, and onychomadesis, and has become a primary HFMD agent in parts of China as EV-A71 declined. Recombination is frequent in the non-structural P2/P3 regions while VP1 stays clonal (Wang et al., Virol Sin 2022), so a whole-genome placement can disagree with VP1 typing for a recombinant strain -- the reason VP1 remains the typing axis. Non-blocking issues🟡 F1. The reference role is described inconsistently across the README and PR [click to expand]The dataset uses two sequences for two purposes: the inferred ancestor aligns and roots the tree, and Kowalik is the default reference for mutation calling in Nextclade Web. The configuration states this unambiguously (
The Reference-types sentence is inherited verbatim from Suggestions:
🟡 F2. The clade scheme is uncited and the documented resolution is imprecise [click to expand]The README advertises "subgenotype assignment" (README.md#L16) and states that naming "follows conventions established in published studies" without naming any. Three problems follow:
Suggestions:
🟡 F3. Two user-facing links are broken or non-resolving [click to expand]Both links are rendered in the Nextclade Web dataset page:
Suggestions:
🟡 F4. Sparse clade G coverage makes genuine divergent G sequences fail QC [click to expand]Clade G has only 5 tips, and the two clade-G example sequences that are not already tree leaves both come back
The clade call itself is correct; the tree simply holds almost no African or environmental G diversity (G is 5 tips, and only 13 of 498 tips are African). Unlike the five deliberately hard cases, these two are not listed under the Suggestions:
🟡 F5. Unsequenced regions of the VP1-only clades appear as deletions in the reference tree [click to expand]Clades E and H are seeded from partial VP1 records. In The impact is limited to tree display, not analysis. A live CLI run confirms queries are unaffected: every example returns 0-9 total deletions, and the E-assigned example returns 0. Nextclade computes query mutations against Suggestions:
🔵 F6. Annotation sequence identifiers disagree [click to expand]One reference sequence carries three identifiers: Suggestions:
🔵 F7. Collection Date is a 261-value categorical coloring [click to expand]
Suggestions:
🔵 F8. CHANGELOG grammar [click to expand]CHANGELOG.md#L3 reads "Initial release of an Coxsackievirus A10 dataset". The changelog is shown in Nextclade Web. Suggestions:
🔵 F9. tree.json has no trailing newline [click to expand]Both the source and generated Suggestions:
Clade distribution8 subgenogroups plus the inferred ancestor, 498 tips [click to expand]
Sampling is heavily weighted toward China (334/498) and toward 2018-2019, thinning after 2022; dated tips span 2008-2025. The balance is defensible against the literature:
The real gaps are clade G (F4) and recent European D diversity: the earliest D tip is 2014, so the outbreak lineages that defined D are absent. Validation summaryStructural checks and official Docker run [click to expand]GFF3 annotation
Reference
Tree
Generated output and registration
Nextclade CLI run (
NotesClick to expand
|
Description of proposed changes
This PR introduces a new Coxsackievirus A10 dataset rooted on the Kowalik reference sequence. It includes all essential files for Nextclade compatibility, such as the reference genome, genome annotation, and detailed documentation. The dataset supports subgenotype assignment, phylogenetic placement, and sequence quality control for Coxsackievirus A10.
New Coxsackievirus A10 Dataset
Dataset documentation:
README.mdexplaining dataset scope, subgenogroup definitions, reference types, and usage instructions.CHANGELOG.mdwith an initial release entry.Reference data and genome annotation:
reference.fasta.Checklist