Skip to content

hcv: initial import - #479

Open
rneher wants to merge 7 commits into
masterfrom
hcv
Open

rneher wants to merge 7 commits into
masterfrom
hcv

Conversation

@rneher

@rneher rneher commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

@rneher
rneher had a problem deploying to refs/pull/479/merge September 27, 2026 17:23 — with GitHub Actions Failure
@ivan-aksamentov

ivan-aksamentov commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Note

✨ Claude Opus 5.5 (with help from GPT-5.6 Sol) is responding on behalf of Ivan

First review of PR #479, adding the nextstrain/hcv/AF009606 dataset, at head 79b33b25 against base cc60ead2.

Testing

Observed

What the dataset is, and where it diverges from its description [click to expand]

The PR adds a whole-genome HCV dataset for all genotypes (1-8). The dataset has these properties:

  • Reference: the H77 RefSeq NC_038882.1 (9,599 nt)
  • Annotation: ten mature-peptide CDS, from core to NS5B
  • Tree: 182 tips, labeled with genotype (clade) and subtype (a custom subtype column)
  • Suggestion fingerprint: a 7-sequence minimizer fingerprint (additional_references.fasta)

The dataset is built by the new nextstrain/hcv workflow. The PR body contains only a preview link. The four source commits are the initial import, a GFF phase-column fix, defaultCds: NS3, and stricter gap penalties. The PR also registers the dataset in collection.json and adds the generated output.

Discrepancies between the description and the data:

  • Contradiction: the dataset path and CHANGELOG.md name AF009606.1, but reference.fasta, pathogen.json and the GFF use NC_038882.1, which is a different H77 clone (F1)
  • Contradiction: the README says that gap-rich regions are excluded from placement. The masks are in a field that Nextclade does not read (F4)
  • Contradiction: the README says that sequences of rare subtypes get the genotype but possibly no subtype. In tests, some get a wrong subtype (F2), and a genotype 8 genome gets no genotype (F3)
  • Diff noise: the PR file list includes data_output/nextstrain/ndv/class-1/.... These files come from the older merge base (5c40b066). Compared with the current base, the PR does not change NDV output (N1)

Background

HCV taxonomy, H77 reference, genome organization, genotype classification [click to expand]

Taxonomy and reference strain

HCV is a positive-sense single-stranded RNA virus (Baltimore group IV). ICTV MSL #41 (ratified 2026) renamed the species to Orthohepacivirus hominis, in genus Orthohepacivirus and the new family Hepaciviridae (ICTV taxon history, Orthohepacivirus). The previous name was Hepacivirus hominis in Flaviviridae (MSL #38, Postler et al., Arch Virol 2023).

H77 comes from a US patient who was infected in 1977 (Ogata et al., PNAS 1991). Two different infectious cDNA clones of H77 are in GenBank:

Accession RefSeq Length Clone / publication
AF009606 NC_004102 9,646 Kolykhalov et al., Science 1997
AF011751 NC_038882 9,599 pCV-H77C, Yanagi et al., PNAS 1997

AF009606 is the standard coordinate reference for HCV numbering (Kuiken et al., Hepatology 2006; used as such in Smith et al., Hepatology 2014). A Nextclade run of AF009606 against this dataset shows how the two clones differ:

  • Coding region: coordinates are identical
  • Substitutions: 32 nt substitutions, 4 of them amino-acid changes (E2:N8S, NS3:G176E, NS4B:H31Q, p7:L44F)
  • Length difference: all of it is in the 3' UTR poly(U/UC) tract. Both clones carry the complete 98-nt 3'X element (Kolykhalov et al., J Virol 1996)

Genome organization

The 9.6 kb genome encodes one polyprotein that is cleaved into core, E1, E2, p7, NS2, NS3, NS4A, NS4B, NS5A and NS5B. The dataset annotates these ten mature peptides. Only core starts with ATG, and NS5B is followed by the polyprotein stop codon (TGA), which is expected for mature peptides.

Genotype / subtype classification

ICTV currently recognizes 8 genotypes and 93 confirmed subtypes. Table 1, dated January 2026, gives one reference genome per subtype. ICTV applies these rules (ICTV Hepacivirus; Smith et al. 2014):

  • New subtype: requires complete coding sequences from at least three epidemiologically unlinked individuals
  • Subtype divergence: subtypes differ by at least 15% of nucleotide positions
  • Recombinants: circulating recombinant forms such as RF2k/1b are listed separately (Table 4)

Genotype 8 was described in 2018 from four patients from Punjab, India (Borgia et al., J Infect Dis 2018). Only subtype 8a is confirmed. Further genotype 8 genomes from the UK are in GenBank (PP092205, submitter title "two new subtypes of hepatitis C virus genotype 8").

Epidemiology and why subtype matters

WHO estimates 47 million people with chronic HCV in 2024 (WHO fact sheet). Genotype 1 is about 44-46% of infections and genotype 3 is 25-30% (Messina et al., Hepatology 2015; Polaris Observatory 2017).

With pan-genotypic antivirals, subtype still matters. Some "unusual" subtypes (1l, 3b, 3g, 4r, 6u, 6v) carry NS5A polymorphisms that reduce susceptibility to direct-acting antivirals (Nguyen et al., J Hepatol 2020). Unusual genotype 1 subtypes are over-represented among treatment failures (Vo-Quang et al., Hepatology 2023). A wrong subtype call is therefore more harmful than no call (F2).

These tools also assign HCV genotype and subtype:

Blocking issues

🔴 F1. Dataset path and CHANGELOG name AF009606, but the reference is NC_038882.1 (= AF011751) [click to expand]

The dataset path nextstrain/hcv/AF009606 and CHANGELOG.md#L4 ("Reference FASTA AF009606.1") refer to one H77 clone. The dataset files use another:

  • reference.fasta: NC_038882.1
  • pathogen.json: NC_038882.1 in attributes["reference accession"]
  • annotation.gff: NC_038882.1 as the sequence ID
  • tree.json: NC_038882.1 as the root sequence

NCBI states that NC_038882 "is identical to AF011751". This is the pCV-H77C clone (9,599 nt). AF009606 is the Kolykhalov & Rice clone (9,646 nt, RefSeq NC_004102). The two sequences differ by 32 substitutions and by 3' UTR indels (see Background). The upstream workflow also uses NC_038882.1 (nextclade/defaults/reference.fasta).

Dataset paths are immutable after release (curation guide). If this ships, the identifier permanently names an accession that the dataset does not use. The dataset list would show reference accession: NC_038882.1 under the path .../AF009606, and users who compare coordinates with AF009606 would get small differences at 32 positions and in the 3' UTR.

Suggestions (pick one):

  • Keep the path and switch the reference to AF009606.1 (or its RefSeq NC_004102.1). This is the community numbering standard (Kuiken et al. 2006), and AF009606 is already a tip in the tree. Coding coordinates do not change, so the GFF CDS ranges stay valid. The sequence ID, the 3' UTR length and the tree root must be regenerated
  • Keep the reference and rename the directory to NC_038882 (or AF011751), and fix the CHANGELOG line
  • Either way, decide now whether a future genotype-specific dataset needs room in the path, for example hcv/all/<accession> next to a later hcv/3a/<accession>. Siblings use both styles (dengue/all, ndv/class-1/AB524405)
🔴 F2. Subtypes missing from the tree are often called as a wrong subtype [click to expand]

The tree contains 46 of the 93 confirmed subtypes. The README says that sequences of unrepresented subtypes "are assigned the genotype, but may not get a subtype assignment". To test this, I ran the reference genome of every confirmed subtype from ICTV Table 1 that NCBI returned (92 genomes) through the dataset with Nextclade 3.23.0:

ICTV reference genome Genotype correct Subtype correct No subtype Wrong subtype
Tip in the tree (32) 32 28 3 1
Not in the tree (60) 60 16 32 12

The genotype is always right. The subtype is wrong for 12 of the 60 genomes that are not in the tree:

Accession ICTV subtype Called QC
KJ439775 1n 1a bad
KJ439778 1m 1a bad
KJ439772 1i 1b mediocre
KJ439768 1d 1b bad
KC248194 1e 1g bad
PQ899568 5b 5a bad
EU246940 6u 6e bad
KJ567651 6xc 6e bad
JX183552 6xb 6k bad
JX183557 6xe 6n mediocre
MH492360 6xg 6n bad
JX183549 6xi 6k bad

The tree labels also disagree with ICTV for three reference tips:

  • DQ278891: labeled 6k from NCBI metadata. ICTV lists it as the 6xj reference, so every query placed near it gets 6k
  • MH590698: 8a in ICTV, no subtype in the tree
  • EF108306 / NC_030791: 7a in ICTV, no subtype in the tree
  • EU246939 (6t) also has no subtype

The cause is that the subtype of the nearest labeled clade is inherited, and clade labels come from NCBI metadata. The PR does not change the privateMutations QC settings. They still mark most wrong calls as "bad" (9/12), but the subtype column itself is wrong. Two of the wrong calls affect clinically relevant unusual subtypes (1-non-a/b, 6u; see Background).

Suggestions:

  • Always include the ICTV Table 1 reference genome of each confirmed subtype in the tree (for example as a forced include in the workflow's include.txt), and label those tips with the ICTV subtype. Let ICTV labels override NCBI labels (via the workflow's genotype_annotations.tsv)
  • Update the README sentence to state the actual behavior and to point users to the QC status for rare subtypes
🔴 F3. A genotype 8 genome gets no genotype [click to expand]

The README says the dataset covers "all known genotypes (1-8)". Genotype 8 has one tip (MH590698, 8a) at the end of a 1,838-mutation branch. The clade_membership label is only on that tip, and the parent nodes have no genotype. A genotype 8 genome from another lineage is placed on the unlabeled backbone:

Query Genotype Subtype Private mutations QC
PP092205.1 (genotype 8, UK) (none) (none) 2,239 bad

Other public genotype 8 genomes are not in the tree, because NCBI has no genotype field for them and the workflow annotates only MH590698 (genotype_annotations.tsv):

Genotype 7 has a similar weakness. It has 3 tips, but EF108306 and NC_030791 are the same isolate (F11), so only two distinct genomes remain. The genotype 7 queries tested here still got genotype 7.

Suggestions:

  • Add the published genotype 8 genomes, with genotype = 8 in genotype_annotations.tsv, and retest PP092205
  • Include the 7b reference (KX092342) and other distinct genotype 7 genomes

Non-blocking issues

🟡 F4. Placement masks are in pathogen.json, where Nextclade does not read them [click to expand]

pathogen.json#L37 defines placementMaskRanges. This key is not in the current input-pathogen-json schema. Nextclade reads placement masks only from tree.json meta.extensions.nextclade.placement_mask_ranges (extensions schema). Community datasets such as neherlab/hiv-1/hxb2 and itps/zikav use the tree field. Because of this:

  • Silent ignore: the masks have no effect on placement
  • False README claim: the README says that gap-rich regions "are excluded from phylogenetic placement", which is not true for this build

I moved the same five ranges into a copy of tree.json and reran 28 genomes. Genotype and subtype calls did not change, and two private-mutation totals changed by 1. So the practical effect today is small, but the configuration and the documentation are wrong.

Suggestions:

  • Export the ranges into the tree:
{
  "meta": {
    "extensions": {
      "nextclade": {
        "placement_mask_ranges": [
          { "begin": 2045, "end": 2070 },
          { "begin": 7035, "end": 7060 },
          { "begin": 7380, "end": 7490 },
          { "begin": 7560, "end": 7590 },
          { "begin": 9375, "end": 9599 }
        ]
      }
    }
  }
}
  • Remove placementMaskRanges from pathogen.json
🟡 F5. Genotype 1 is missing from the dataset-suggestion fingerprint [click to expand]

pathogen.json#L13 sets minimizerIndex.references: ["additional_references.fasta"]. When this list is set, scripts/minimizer uses only these files and does not add the main reference. The file has one genome each for genotypes 2-8 and none for genotype 1. The upstream include.txt lists AF009606 under "Minimizer references (one per genotype)", but config.yml minimizer_references does not contain it.

nextclade sort with the generated data_output/minimizer_index.json gives these results:

Input (27 genomes) Genotype 1 suggested Other genotypes suggested
Full genomes 6/6 (lowest scores: 0.16-0.21) 21/21
NS5B fragment (H77 8256-8644) 0/6 15/21
E2 1/6 18/21
Core/E1 (869-1292) 0/6 12/21

Genotype 1 is the most common genotype (see Background). Nextclade Web does not suggest this dataset for most genotype 1 partial sequences.

Suggestions:

  • Add NC_038882 (or AF009606, see F1) and a 1b genome to additional_references.fasta, and add them to minimizer_references in the workflow
🟡 F6. NS5A V3 alignment artifacts trigger frameshift QC [click to expand]

Frameshift QC is "mediocre" for 5/20 example sequences (6a x3, 2b x2) and for 2/7 additional references (3a, 7). All flagged frameshifts are in NS5A codons 376-408, which is variable region V3 (see Background). Each one is a compensated indel pair, for example:

  • DQ480521: 1-nt deletion at 7391 and 16-nt insertion at 7480 (net in-frame)
  • AF238486: 10-nt and 2-nt insertions (net in-frame)

These are alignment artifacts against a genotype 1a reference, not real frameshifts. Amino acid changes that Nextclade reports in this NS5A window for non-1 genotypes are unreliable. None of these changes removed the artifacts:

  • alignmentPreset: "high-diversity"
  • penaltyGapOpenOutOfFrame from 31 up to 90

ignoredFrameShifts needs an exact codon-range match (qc_rule_frame_shifts.rs), so it cannot cover the variable ranges. Overall QC stays "good" because scoreWeight: 40 alone does not reach "mediocre".

Suggestions:

  • State in the README that frameshifts and amino acid changes in NS5A codons ~370-410 are expected alignment artifacts for non-1 genotypes
  • Optionally list the recurrent exact ranges in ignoredFrameShifts (for example NS5A 381-388 for 2b, 397-407/397-408 for 6a), or reduce frameShifts.scoreWeight
🟡 F7. README has unfilled TODO placeholders [click to expand]

Nextclade Web shows the README to users. It contains three placeholders:

Suggestions:

  • Workflow: https://github.com/nextstrain/hcv. Path: the final dataset path (F1)
  • Partial-sequence text: use the measurements in F8
🟡 F8. Core/E1 typing fragments often fail to align [click to expand]

I cut fragments from the 27 examples and additional references (reference coordinates from the Nextclade alignment) and ran them through the dataset:

Fragment (H77 coordinates) Aligned Genotype and subtype = full-genome call
Core (342-914) 27/27 27/27
E2 (1491-2579) 27/27 27/27
NS5B complete (7602-9374) 27/27 27/27
NS5B typing fragment (8256-8644) 27/27 27/27
Core/E1 typing fragment (869-1292) 18/27 18/18

Core/E1 is one of the two standard typing regions (Murphy et al. 2007). For 9 genomes (2b x3, 3a x4, 5a, 6t), the fragment fails with "seed alignment was unable to find any matches". alignmentPreset: "high-diversity" gives the same result. Users with core/E1 amplicons from these genotypes get an error instead of a result.

Suggestions:

  • Document the behavior in the README (F7)
  • Optionally tune seed parameters (kmerLength, allowedMismatches, minMatchLength) against core/E1 fragments of all genotypes
🔵 F9. CHANGELOG has no trailing newline [click to expand]

CHANGELOG.md has no newline at the end of the file. The accession on its last line is part of F1.

Suggestions:

  • Add a trailing newline
🔵 F10. No compatibility block [click to expand]

pathogen.json has no compatibility field. Sibling datasets in the nextstrain collection set {"cli": "3.0.0-alpha.0", "web": "3.0.0-alpha.0"}. Without it, older Nextclade versions list the dataset even when they cannot use it.

Suggestions:

  • Add a compatibility block that matches the siblings
🔵 F11. The tree has duplicate RefSeq/GenBank tips [click to expand]

Five tip pairs are the same record under a RefSeq and a GenBank accession (zero-length sister branches):

  • NC_009826 / Y13184
  • NC_009827 / D84262
  • NC_030791 / EF108306
  • NC_004102 / AF009606
  • NC_009825 / Y11604

These pairs inflate the tip counts. For genotype 7, 3 tips are only 2 distinct genomes.

Suggestions:

  • Exclude the RefSeq copies (NC_009823-NC_009827, NC_004102, NC_030791) in exclude.txt and keep the GenBank originals. Keep the dataset reference tip
🔵 F12. Tree defines a coloring with no data [click to expand]

tree.json meta.colorings includes gt, but no node has a gt attribute. The Auspice/Nextclade tree view shows an empty color-by option.

Suggestions:

  • Remove gt from auspice_config.json in the workflow, or populate it

Clade distribution

8 genotypes, 46 subtypes, 182 tips [click to expand]
Genotype Tips % Subtypes (tips without subtype)
1 27 14.8 1a, 1b, 1c, 1g
2 38 20.9 2a, 2b, 2c, 2f, 2i, 2k, 2m
3 22 12.1 3a, 3b, 3g, 3i, 3k
4 34 18.7 4a, 4d, 4f, 4g, 4l, 4m, 4n, 4o, 4r, 4v (1)
5 4 2.2 5a
6 53 29.1 6a-6r, 6v (5)
7 3 1.6 (3)
8 1 0.5 (1)
Total 182 46 subtypes, 10 tips without subtype
  • Sampling: the workflow samples up to 10 genomes per genotype/subtype. As a result, genotype 6 (29%) outweighs genotype 1 (15%), although genotype 1 causes about 45% of infections
  • Sparse genotypes: genotypes 5, 7 and 8 have fewer than 5 tips (F3)
  • Dates: 67 of 182 tips have a date, from 1997 to 2017. NCBI has about 7,000 HCV records of 8.5-10 kb

Validation summary

Validation checks [click to expand]

Reference

  • 9,599 nt, no ambiguous bases, header NC_038882.1 matches attributes["reference accession"] and is identical to NCBI NC_038882.1
  • Fail: path and CHANGELOG name AF009606 (see F1)

GFF3 annotation

  • 10/10 CDS lengths divisible by 3, mature-peptide boundaries contiguous (342-9374), TGA stop after NS5B, phase column 0
  • CDS coordinates identical to tree.json meta.genome_annotations; defaultCds: NS3 exists

pathogen.json

  • Schema URL and schemaVersion valid; files entries match the directory
  • Fail: placementMaskRanges is not a pathogen.json field (see F4)
  • Fail: minimizer fingerprint lacks genotype 1 (see F5)

Tree

  • Auspice v2, 182 unique tips; root_sequence.nuc equals reference.fasta
  • Reference NC_038882 is a tip; reference run gives 0 mutations
  • Fail: labels disagree with ICTV for 3 reference tips (see F2)

Examples

  • 20 sequences; 0/20 are tree tips, so all test runtime placement
  • The 7 additional references are all tree tips

Generated output

  • README, CHANGELOG, GFF, reference and examples are byte-identical to data_output/.../unreleased/
  • tree.json is semantically identical (reformatted); output pathogen.json only adds version.tag
  • index.json and minimizer_index.json: compared with the current base, only the HCV entry is added
  • Registered in data/nextstrain/collection.json
  • Fail: CHANGELOG has no trailing newline (see F9)

Nextclade CLI run (3.23.0, docker)

  • Reference: genotype 1, subtype 1a, 0 mutations, QC "good"
  • AF009606: 1a, 32 substitutions, 0 private mutations, QC "good"
  • Examples: 20/20 overall "good"; frameshift QC "mediocre" in 5/20 (see F6); mixed sites "mediocre" in 3/20; private mutations 0-737
  • Additional references: 5/7 "good", EF108306 "mediocre", EU246939 "bad" (22 mixed sites)
  • ICTV Table 1 reference genomes (92): genotype 92/92; subtype 44 correct, 35 none, 13 wrong (see F2)
  • Genotype 8 PP092205.1: no genotype, QC "bad" (see F3)
  • Partial fragments: see F8; dataset suggestion: see F5

Notes

Click to expand
  • N1. The NDV output files in the PR file list (data_output/nextstrain/ndv/class-1/...) come from the older merge base 5c40b066. Master's own rebuild cc60ead2 produced the same files, so merging changes nothing for NDV. A rebase onto master removes them from the diff
  • N2. The upstream workflow nextstrain/hcv is out of sync with this PR. A regeneration from the workflow would undo these PR commits:
    • N2.1. alignmentParams: the workflow has gap penalties 10/16/18, kmerDistance: 50, allowedMismatches: 8, minSeedCover: 0.1
    • N2.2. mixedSitesThreshold: the workflow has 200, the PR has 20
    • N2.3. defaultCds: missing in the workflow
    • N2.4. GFF phase: the workflow GFF has ., the PR has 0
    • N2.5. config.yml says the mask ranges "should match placementMaskRanges in the pathogen.json" (see F4)
  • N3. mixedSitesThreshold: 20 marks EU246939, which is a tree tip and minimizer reference, as "bad". HCV consensus sequences often carry real ambiguous bases from within-host diversity (Poon et al., PLoS Pathog 2007). A value between 20 and the workflow's 200 may fit better
  • N4. The privateMutations settings (typical 500, cutoff 1000) fit well. Examples of represented subtypes stay "good" (up to 737 private mutations), and most wrong subtype calls in F2 are "bad"
  • N5. snpClusters is disabled, which is appropriate for a dataset with 30% between-genotype divergence
  • N6. Examples cover genotypes 1-4 and 6 only, because all public genotype 5, 7 and 8 genomes are already tree tips. After F3, a genotype 8 example (for example PP092205) would test the difficult case
  • N7. Recombinants such as RF2k/1b are listed in ICTV Table 4. The workflow drops NCBI records labeled "recombinant", and the README says that recombinants get the label of their closest relative. A mixed-genotype QC signal or a note on breakpoint 3186 would help users with 2k/1b sequences
  • N8. Consider "experimental": true until F2 and F3 are addressed, because the README scope claims full genotype and subtype coverage

@ivan-aksamentov

ivan-aksamentov commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Probably too early, but I had some spare tokens before reset, so decided to run my AI review

This branch had an error being deployed

1 failed (outdated) deployment
refs/pull/479/merge — 61b0bf59 Deployed Sep 27, 2026 by rneher via build-and-deploy-datasets #1537
refs/heads/hcv — 61b0bf59 Deployed Sep 27, 2026 by rneher via build-and-deploy-datasets #1536
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants