综述文章 | 拷贝数变异(CNV)检测、疾病关联及其扩展的计算策略

 

摘要

        拷贝数变异(copy number variations, CNVs)是构成人类遗传多样性、进化及疾病易感性的关键结构变异。测序技术与计算方法的进步提高了CNV检测能力,然而关联研究仍受方法学局限和缺乏标准化所困扰。本综述概述了胚系CNV检测与疾病关联的计算策略。我们强调CNV分析在揭示复杂性状与疾病风险的遗传贡献方面的价值,并勾勒出一套包含关键基准测试(benchmarking)方法在内的分析工作流程。我们还讨论了推进CNV检测与关联分析所面临的当前挑战和未来方向。

 

引言

        结构变异(structural variations, SVs)是至关重要的人类遗传变异,通常定义为涉及50个或更多DNA核苷酸的改变(1)。这些SV涵盖广泛的基因组变化,包括拷贝数变异(CNVs)、插入(≥50 bp,包括移动元件插入,mobile element insertions, MEIs)、倒位(inversions)以及染色体异常(2, 3)(图1A)。研究SV对于理解人类遗传多样性、其在进化中的作用及其功能效应至关重要(综述见 Collins 和 Talkowski(2))。

        CNV是一类以特定DNA区段拷贝数改变为特征的SV,是研究最频繁、流行度最高的SV类型之一。既往研究表明,插入与CNV合计占人类基因组中SV的绝大多数,常超过所有检出SV的90%(2)。然而,长读长测序(long-read sequencing, LR-seq)研究显示,SV类别的相对分布可随数据集而异,插入与缺失的检出频率往往相当。因此,观察到的SV谱受测序技术、分析方法和基因组背景的影响(4)。

        CNV涉及基因组内的缺失或重复,通常定义为跨越50个碱基对或以上、且常超过1 kb的变异。这些变异相当普遍,每个基因组约出现10,000处,其中约800处大于10 kb(2)。这些结构改变构成遗传多样性的重要部分,约涵盖人类基因组的5%–10%(5, 6)。

        一个基因座上的CNV可作为复发型或非复发型事件出现。复发型CNV通常由低拷贝重复等基因组特征介导,导致不同个体间断点相似;而非复发型CNV则表现出可变的断点和结构(图1B)。除结构特征外,CNV还发挥多种功能效应,这些效应对下游解读很重要,但并不直接属于计算检测工作流程的一部分。

        图1. 拷贝数变异(CNVs)的结构与功能特征。(A)结构变异(SVs)是基因组改变,包括缺失(DEL)、重复(DUP)、插入(INS;≥50 bp,含移动元件插入,MEIs)、倒位(INV)、复杂重排(CPX)和染色体异常(CA,包括相互易位和非整倍体)。这些包括平衡型(BA)和非平衡型(UB)。平衡型SV保持DNA剂量但可能破坏基因结构,而非平衡型SV导致DNA获得或丢失。CNV是SV中通常分为降低拷贝数的DEL和增加拷贝数的DUP的类型。此外,还可出现复杂的CNV结构,如DUP-三倍化(TRP)/INV-DUP构型。(B)复发型与非复发型CNV的结构和断点模式。参考基因组(REF)被划分为区段,关键区域内含有剂量敏感基因。低拷贝重复(low-copy repeats, LCRs)介导非等位同源重组(non-allelic homologous recombination, NAHR),导致复发型CNV。串联旁系同源重复(tandem paralogous repeats, TPRs)也可介导这些重排。CNV以彩色矩形表示(红色为DEL,蓝色为DUP,橙色为倒位TRP结构)。区段1包含两侧翼有LCR的关键区域基因,使其在不同个体间易产生大小和内容相同的复发型CNV。区段2显示在TPR内可能发生复发型重排,并改变位于重复序列内部的剂量敏感基因的拷贝数。区段3展示非复发型DUP-TRP/INV-DUP结构。区段4代表非复发型重排,其在旁系同源重复内显示断点聚集,与复发型病例中观察到的紧密聚集形成对比。区段5展示无任何断点聚集或成组的非复发型重排,导致不同个体间出现完全独特的CNV。最小重叠区(smallest region of overlap, SRO)指不同个体间重叠CNV中被一致改变的基因组最小区段。图B的构思参考了(7)中描述的复发型与非复发型CNV示意图。

        如图1C所总结,主要机制之一是基因剂量改变:缺失可能导致单倍剂量不足(haploinsufficiency),重复则可能导致三倍敏感(triplosensitivity),尤其是在剂量敏感基因中(2, 8)。CNV还可破坏基因完整性并改变调控,例如扰乱拓扑关联结构域(topologically associating domains, TADs)以及改变增强子-基因相互作用(3, 8, 9)。此外,CNV可对基因表达产生实质性影响,可作为表达数量性状基因座(expression quantitative trait loci, eQTLs)(10, 11),或在复杂重排背景下促成基因融合的形成(12)。缺失还可能暴露同源染色体上的隐性变异,从而揭示原本保持沉默的致病等位基因(8)。

        图1. 拷贝数变异(CNVs)的结构与功能特征。(C)CNV的功能后果。CNV可通过基因剂量改变、基因结构破坏、调控架构变化(包括拓扑关联结构域)、对基因表达的影响以及基因融合的形成或隐性变异的暴露,影响基因和蛋白功能。这些功能后果对于解读CNV的表型影响至关重要。

        由CNV或基因组重排引起的基因组病(genomic disorders, GDs)为CNV的致病性提供了确凿证据(13)。许多已被充分表征的GD,尤其是由复发型大片段CNV介导的GD,涉及跨越多个基因的兆碱基级区域(14)。虽然其中一些疾病具有单一显性遗传驱动因素,但另一些遵循寡基因和多基因模型,凸显了遗传架构的复杂性以及修饰效应的影响(15-17)。在一般人群中发现的若干复发型GD相关CNV,即使没有明显的临床疾病,也可影响生理性状,如身高(18-20)。除经典GD外,CNV还与广泛的神经发育及精神疾病相关,包括精神分裂症、自闭症和双相情感障碍(14, 21, 22)(图2)。

        总体而言,CNV被广泛认为是疾病易感性的重要贡献因素,在某些情况下其效应量可接近孟德尔疾病中观察到的水平。例如,22q11.2基因座的变异即体现了这种影响:缺失与精神分裂症风险增加约60倍相关,而重复可将风险降低至基线水平的约15%(8, 21)。与此同时,许多CNV外显率不完全且表现度可变,反映了其遗传架构的更广泛复杂性(2)。

        图2. 疾病相关拷贝数变异(CNVs)及其相关表型的全基因组分布。为说明CNV的广泛影响,我们可视化了 Collins 等(14)的数据,该研究显示大片段罕见CNV与全基因组范围内多种人类疾病和性状显著相关。该研究在多个队列中进行了大规模荟萃分析,识别出163个与多种人类疾病和性状显著相关的罕见大片段CNV基因座,尤其是神经发育及精神疾病。Circos图显示1–22号染色体,外圈为染色体模式图(ideograms)。第二轨道显示沿每条染色体的CNV计数密度图。第三轨道描绘各CNV基因座,红色方块代表缺失,蓝色方块代表重复。第四轨道为热图,显示映射到每个基因座的人类表型本体(Human Phenotype Ontology, HPO)条目数(由低到高强度)。最内层连线将基因组基因座连接到HPO类别,按图例所示颜色编码。

        大规模生物样本库(biobanks)对于推动该领域至关重要。通过将全基因组CNV发现与大规模、广泛表征的表型数据相关联,英国生物样本库(UK Biobank)等资源能够在数万名个体中开展系统的CNV关联分析。这些分析可细化外显率估计、识别跨性状的多效性效应,并揭示来自局部和全基因组遗传背景的修饰影响(8, 18, 24-27)。测序技术和计算方法的进步使得更大规模、更高分辨率和更灵敏的CNV检测成为可能,其中包括人类基因组结构变异联盟(Human Genome Structural Variation Consortium)等大规模工作的重大贡献,该联盟利用LR-seq和单倍型解析组装改进了SV的发现与表征(28-32)。然而,CNV关联研究在最佳实践方面仍存在若干空白,挑战涵盖方法学问题以及缺乏标准化和稳健的数据基础设施。这些局限阻碍了准确且可重复的CNV检测与关联分析。一个主要障碍是不同算法和平台之间CNV检出结果的可变性大且共识度低(33-36)。

        工具性能高度依赖具体情境,其优缺点因测序技术、数据类型和分析策略而异。因此,没有任何单一算法能在所有平台和用例中持续优于其他算法(33, 34, 37-41)。现有工具众多且结果各异,选择合适的工具令人无所适从。大多数当前工具比较依赖于工具开发者提供的基准,可能引入偏倚。

        若缺乏独立基准测试,研究人员和临床医生可能做出次优选择,从而影响分析和临床解读(33, 37-40, 42-45)。此外,与基于单核苷酸多态性(single nucleotide polymorphism, SNP)的关联分析(已有PLINK(46, 47)和REGENIE(48)等标准化工具)不同,目前尚无被普遍采用或作为社区标准的CNV关联分析工具(8, 49),尽管正在进行的社区工作正在积极开发和基准测试该领域的方法。现有工具往往缺乏用户友好性、必要的报告功能,或未能整合必要的分析步骤(49-52)。

        在本综述中,我们聚焦于胚系CNV,因为与通常在癌症基因组学和嵌合(mosaic)背景下研究的体细胞CNV相比,胚系CNV的检测、解读和关联分析依赖于不同的研究设计和计算框架。虽然体细胞CNV在癌症及非癌症疾病中均具有公认的作用(40, 53-65),但其分析涉及根本不同的方法学考量,因此不在本综述范围内。据此,我们专注于胚系CNV检测与关联分析的计算方法。

        我们首先强调CNV关联分析在揭示SNP分析常遗漏的复杂性状遗传贡献方面的价值。随后我们提供CNV关联分析的一般工作流程。接着介绍CNV检测的平台和方法,并概述基准测试框架及比较CNV检测工具所面临的挑战。接下来,我们评述CNV检出与关联分析工具,重点关注独立的基准测试研究,以实现无偏倚的性能评价。

        在缺乏独立证据的情况下,我们以对社区推荐工具的批判性概述作为补充,为研究人员和临床医生选择适当的CNV检测与关联方法提供实用、循证且针对具体情境的建议。最后,我们以当前挑战和CNV检测及关联分析的未来展望作为结论。

 

CNV关联研究的独特贡献

        CNV关联分析提供的洞见超出了基于SNP的方法所能及的范围,为理解复杂性状的遗传架构提供了独特机遇并拓展了我们的认识(8, 20, 26, 27, 66)。在某些情况下,出生体重、总胆固醇、低密度脂蛋白(LDL)胆固醇和载脂蛋白B等性状已被证明与CNV负荷相关,提示存在一个由罕见、高效应CNV或效应温和的较常见变异所塑造的额外遗传架构层面(27)。关于利用CNV建模进行性状关联检验的文献日益增多(附加文件1:图S1,附加文件2:表S1),但大多数仍依赖分辨率有限的SNP芯片(8)。传统SNP芯片缺乏可靠检测小片段CNV所需的分辨率,通常对基因组中存在的全部CNV谱系不敏感,尤其是那些较小但可能具有影响的CNV(67-69)。此外,基因组CNV基因座与全基因组关联研究(GWAS)信号之间存在大量重叠,且其在重组热点中富集,提示CNV可能解释一部分仅凭SNP无法解释的”缺失遗传力”(missing heritability)(70)。

        SNP与CNV信号会相互影响彼此的测量与解读。在基于芯片的基因分型中,SNP检出算法通常假定固定的二倍体拷贝数状态,捕捉潜在CNV变异的能力有限。因此,未被建模的CNV可导致错误的基因型分配以及看似偏离哈迪-温伯格平衡(Hardy-Weinberg equilibrium)的现象。反之,在CNV检出时忽略SNP信息可能遗漏等位基因特异性的获得和丢失,并限制利用CNV与邻近SNP之间连锁不平衡(linkage disequilibrium, LD)的能力(71)。许多CNV(尤其是罕见或结构复杂的CNV)很难被SNP标记(tagging),这意味着标准方法无法检测到它们的影响(8, 72-78)。其中一个主要原因在于,CNV(尤其是复发型CNV)可因局部升高的突变率而在特定基因组基因座反复出现,产生多个功能后果相似的独立事件。CNV的新发(de novo)突变率估计约为每个个体0.2个事件,反映了某些基因组区域易于发生复发性结构改变的倾向。因此,CNV往往与特定单倍型之间缺乏一致关联,使其难以被SNP标记(72-75, 79)。由此,CNV与邻近SNP之间的LD通常较弱或可变,尤其是在重复序列背景下,那里升高的CNV突变率和增加的基因分型错误进一步破坏了LD的稳定性。此外,多等位CNV(multiallelic CNVs, mCNVs)——其拷贝数超出二倍体预期的广泛范围——可打破典型的LD模式,并产生无法被短变异标记的群体特异性”失控单倍型”(runaway haplotypes)。这些mCNV与邻近SNP的平均观察LD往往也较低(2, 72, 78, 80, 81)。某些CNV的稀有性和结构多样性也限制了标准SNP关联方法检测或准确建模它们的统计效力,而复杂的CNV重排则给准确解读和标记带来额外挑战(47, 52, 66)。根据威康信托病例对照联盟(Wellcome Trust Case Control Consortium)的数据,79%的常见CNV(最小等位基因频率(MAF)> 10%)可被高LD的邻近SNP标记。然而,仅有约22%的较不常见CNV(MAF低于5%)能被SNP良好标记(77),而且这种标记在非洲或非裔美国人血统人群中效果更差,因为这些人群中SNP与CNV之间的LD更低(78)。

        综合来看,当考虑全基因组CNV的完整谱系(尤其是罕见、复杂和多等位变异)时,强SNP-CNV标记并不常见。因此,CNV特异性关联研究在揭示SNP关联研究无法检测的遗传信号方面发挥着至关重要的作用。例如,在一项分析精细定位(fine-mapped)CNV关联区域的最新研究中,17%(133个区域中的23个)被归类为”CNV独有”(CNV-only),凸显了CNV分析对遗传发现的独特贡献(66)。CNV独有关联指仅能通过CNV-GWAS而非SNP-GWAS检测到的遗传信号。CNV独有关联的实例可说明CNV-GWAS的附加价值。例如,影响SPDYE1、SPDYE6及POLR相关基因的CNV已与睡眠时型(chronotype)相关联(66)。值得注意的是,这些基因此前从未被英国生物样本库的SNP-GWAS与睡眠时型联系起来,提示这些区域的CNV可能以剂量依赖方式影响个体的睡眠模式(66)。还发现了其他CNV独有关联,例如ZDHHC11B基因区域与用力呼气量(FEV)/用力肺活量(FVC)比值相关,以及CDK11A基因区域与身高相关(66)。

        CNV发挥效应的机制也可能与SNP存在根本差异,因为CNV可通过剂量改变发挥作用,其表型后果可能与SNP模型无法完全捕捉的后果相反或相同(20, 26)。CNV分析还增进了我们对基因功能和剂量敏感性的理解,产生或印证了关于特定基因如何影响生物学过程和疾病风险的假说(20, 26, 27)。通过在一般人群(而非仅临床队列)中研究致病性CNV,研究人员可以更好地刻画从严重早发疾病到轻度或无症状表现的完整表型结局谱。这与可变表现度和不完全外显的模型相一致,提供了对基因型-表型关系更细致的视角(26, 27, 82-84)。

        从临床角度看,CNV关联研究可生成”发病图谱”(morbidity maps),帮助临床医生预测和监测特定CNV携带者的潜在结局(26, 85, 86)。虽然BRCA1和LDLR等基因的功能丧失变异已知可增加疾病风险(与变异类型无关),但CNV关联研究通过刻画这些改变在人群层面的影响提供了互补性检测。特别是,它们能够在大型队列中评估外显率、可变表现度、合并症及总体CNV负荷,从而细化我们对疾病风险和临床结局的理解(26)。这些洞见改善了风险分层,并支持临床决策,尤其是对于复发型或多基因CNV。

 

CNV关联分析的一般工作流程

        CNV关联分析的工作流程识别与人类性状和疾病显著相关的CNV,并解读其生物学和临床意义(图3)。

        图3. 拷贝数变异(CNV)关联分析的工作流程。该流程始于从多种平台获取数据,包括SNP芯片、短读长外显子组测序(ES)、短读长基因组测序(GS)和长读长测序(LR-seq)。此处ES以读长深度(read depth)信号表示,这是其CNV检测的主要输入。在初步归一化和质量控制(QC)之后,使用平台特异性CNV检出器或集成/合并(ensemble/merging)策略进行CNV检出。本图列出的一些检出器(如DRAGEN)代表适用于多种测序模式(包括ES和GS)的综合性端到端分析平台,而非独立的CNV检测算法。值得注意的是,存在两个例外:探针水平分析(直接利用SNP芯片的归一化信号强度)和外显子水平分析(检验ES中每个外显子的归一化读长深度)。这两种例外绕过了CNV检出,直接处理定量信号。在CNV检出(或这些例外中的信号水平方法)之后,实施检出后QC,通过过滤低质量事件和含噪声区域来确保可靠性。随后,检出的CNV以多种分辨率进行分类以进行关联检验,包括探针水平(信号或代理状态)、平铺窗口(tiling-window)、外显子水平、区域水平(CNVR)、CNV水平或基因水平。尽管理论上可以设想其他分组方案,但这些水平代表了实践中采用最频繁的方法。统计关联检验使用标准回归模型(线性、逻辑、混合模型)、负荷检验(burden tests)或基于核的方法进行,并得到日益增多的专用CNV关联工具的支持,如PLINK(46, 47)、ParseCNV2(49)、CNVtools(87)、CNVRanger(88)和CNest(66)。效应建模框架(基于剂量、分类、概率或类型特异性的)以及关联后分析步骤(精细定位、重复验证、正交验证、与SNP-GWAS整合、通路分析和临床注释)进一步细化生物学和临床推断。

        图4. CNV信号表示与CNV区域(CNVR)定义策略。(A)展示了各平台CNV信号表示的示意图。对于SNP芯片,呈现归一化指标,如B等位基因频率(BAF)和对数R比值(LRR),以及拷贝数状态。在基于短读长测序的方法中,相关的CNV证据包括读长深度(RD)、不一致读长对(RP)、分裂读长(SR)、组装(AS)和长读长测序(LR-seq),采用基于比对或基于组装的方法,通过比对缺口或辅助比对识别缺失和重复。这些信号提供互补的CNV证据:RD反映剂量变化,RP提示异常的插入片段大小或方向,SR可实现断点解析;这些相同的特征也用于人工检查(例如在IGV中),以区分真实变异与技术假象。关于检测方法和信号处理的更多技术细节见附加文件3:注释S1(2, 28, 29, 57-61, 89-104)和附加文件3:注释S2(28, 95, 104-133)。(B)CNVR定义及病例-对照分析中常见CNVR模式的示意图。不同个体间重叠的CNV检出可通过多种方法合并为CNVR,包括裁剪(trimming)、互惠重叠(RO)和基于片段的技术(50)。鉴于CNV大小和断点的可变性,病例-对照数据中可能出现不同的CNVR模式。行为良好的CNVR在不同样本间显示高度一致的CNV边界。随机边界变异反映断点中的轻微随机波动。多个显著CNVR代表重叠CNV的独立簇。CNV半岛(CNV peninsula)指与共享区域相邻的、探针支持有限的小型病例特异性延伸。中心共识伴可变延伸指具有可变侧翼边界的共享核心CNV。对照侵占(control encroachment)发生在对照中的CNV与主要由病例定义的区域重叠时(52)。

1

数据获取与初始处理

        工作流程始于从基于芯片或基于测序的平台获取数据。关于CNV检测平台和方法学的详细概述见附加文件3:注释S1。对于基于芯片的数据集,原始信号经处理以获得对数R比值(LRR,代表相对信号强度)和B等位基因频率(BAF,显示每个基因座的B等位基因比例)。对于基于测序的数据集,提取读长深度谱和比对特征,根据CNV检出算法的不同可能纳入分裂读长或不一致读长对。此阶段包括归一化和QC,以在CNV检测前减少技术假象(4, 19, 20, 24, 26, 66, 86, 135-142)。

2

CNV检出

        预处理后,应用CNV检出算法推断个体拷贝数状态(关于CNV检测策略和分段方法的详细信息见附加文件3:注释S2和附加文件3:注释S3(104, 114, 121, 143-157))。CNV检测性能强烈依赖于底层平台,在分辨率、断点准确性和对CNV大小的敏感性方面存在关键差异。不同工具采用统计模型整合各种信号特征来检出CNV(附加文件1:表S2(66, 106-108, 110-115, 120-123, 126-128, 130, 132, 143, 145, 146, 148, 150-152, 158-208)和附加文件1:图S2-S8(209))。为提高可靠性,通常并行使用多种算法(28)。

通过集成与合并策略进行CNV检测的方法与工具

        用于CNV检测的集成(ensemble)与合并(merging)方法旨在克服单一检出工具的局限性,如高假阳性率、低一致性和检出不完整(28)。由于没有任何单一CNV检测方法能在各种情境下始终提供最佳结果,因此结合多种工具的优势已成为一种有前景的CNV检测策略。

        这些方法可根据其合并与细化变异检出结果的主要策略进行大致分类。这些集成与合并策略已在不同平台上开发,包括SNP芯片、外显子组测序(ES)、基因组测序(GS)和LR-seq,其适用性通常具有平台特异性。一类是简单共识或基于重叠的方法,依赖变异检出的直接重叠,或要求最少数量检出器对某一变异达成一致以定义高置信度检出。虽然直接,但这些方法未必总能达到最高精度或召回率。例如,为GS数据开发的HugeSeq将两个或更多算法检出、互惠重叠至少50%的SV和CNV定义为高置信度变异,使用BEDtools整合BreakDancer、Pindel、CNVnator和BreakSeq等工具的输出(184)。NextSV专为从低覆盖度LR-seq数据检测SV而设计,整合多个SV检出器的结果以提高稳健性。它产生两个检出集:灵敏集(所有检出的并集,最大化召回率)和严格集(检出的交集,优先保证精度)。例如,对于缺失,检出根据互惠重叠阈值(通常为50%或更高)进行合并(174)。SURVIVOR是另一个广泛使用的工具,根据用户定义的标准(如坐标距离、SV类型和染色体)过滤和合并SV,支持Delly、LUMPY、Pindel和cn.MOPs等工具(177)。

        虽然combiSV也整合多个检出器的输出,但它并非简单的基于重叠的方法。它是一个专为通过整合多个SV检测器输出来改进LR-seq数据SV检测而设计的新近集成工具。它接收多达六个不同工具(即cuteSV、pbsv、Sniffles、NanoVar、NanoSV和SVIM)的变异调用格式(VCF)输出,并整合其检出结果以产生具有更高召回率和精度的共识SV检出集。与简单的基于重叠的方法不同,combiSV使用检出器特异性优先级排序,基于基准测试结果为每个SV参数(如位置、长度和基因型)选择最准确的信息,属于高级集成方法类别。它还允许用户设置最小支持检出器数量和每个变异的读长覆盖度阈值(172)。

        更复杂的一类涉及机器学习和统计整合方法,它们使用数据驱动模型为变异检出分配权重或整合特征,从而获得更好的准确性和更少的假阳性。为ES数据设计的CN-Learn是一个基于随机森林的框架,使用GC含量和可映射性(mappability)等特征,在经验证的CNV子集上训练,整合来自CANOES、CODEX、XHMM和CLAMMS等多种算法的CNV检出(194)。针对GS数据的FusorSV是结构变异引擎(Structural Variation Engine)的一部分,使用数据挖掘方法针对真实集(truth set)评估工具性能。它采用互斥策略整合检出结果,在最大化灵敏度的同时最小化假阳性(210)。

        另一组工具属于带细化或验证流程的高级合并类别,它们不仅整合检出结果,还细化断点、重新进行基因分型,并使用额外证据进行验证。例如,EnsembleCNV使用两阶段流程从SNP芯片数据中检测和分型CNV,包括初始检测和利用局部似然模型细化边界的重新分型步骤(132)。针对GS数据优化的MetaSV使用工具内和工具间合并策略整合多个SV检测工具的检出,通过SPAdes整合局部组装、细化断点、进行基因分型并注释变异(190)。Parliament2专为GS数据的可扩展分析而设计,整合多个SV检出器,使用SURVIVOR进行合并、SVTyper进行验证,并根据支持证据为每个SV分配质量评分(206)。同样为GS数据开发的SVMerge采用模块化方法,整合来自各种工具的SV检出并使用局部从头组装(de novo assembly)进行细化,使其可扩展至未来的SV检出工具(202)。

        集成与合并算法改进了CNV检测,但面临关键局限性,包括缺乏标准化、测序覆盖度可变、依赖特定算法组合、基准测试不一致以及独立工具尚不成熟。关于SV检测集成算法的概述,读者可参阅Ho等的详细综述(28)。

基准测试概述

        对SV(包括CNV)进行基准测试是基因组研究中的关键过程,旨在评估和验证变异检出方法的准确性。基因组测序技术的不断进步增强了对SV的检测,尤其是随着LR-seq的出现,它可以对这些变异进行更精细的表征(211)。

        基准测试包含四个核心组成部分。第一是建立真实数据(211),即金标准或真实参考数据集(如”瓶中基因组”联盟(Genome in a Bottle Consortium, GIAB)生成的数据集(212)),分析结果可与之比较。这些真实集(也称为高置信度或真实参考检出集)通常通过整合多种测序技术、变异检出方法和广泛的人工整理生成,以最小化平台特异性偏倚(211)。然而,GIAB联盟等广泛使用的资源仅来源于有限的个体,可能无法完全反映跨人群的临床相关CNV的多样性或谱系(211)。用于评估SV(包括CNV)的基准数据集有多种可用(由Majidian等综述(211))。GIAB衍生真实集的一个重要补充是Platinum Pedigree资源(213),它利用多代家系(CEPH-1463)中的孟德尔遗传,该家系使用PacBio HiFi、Illumina和Oxford Nanopore Technologies测序。通过追踪家系内的单倍型传递,该资源在全基因组范围内验证变异检出,包括单核苷酸变异、插入缺失(indels)、串联重复和SV,覆盖常规真实集常排除的困难基因组区域,优先保证灵敏度和完整性而非特异性。第二,涉及用于将分析结果与真实数据比对和比较的特定工具和算法。第三,使用性能指标,通过标准统计度量全面评估一致性和诊断准确性。这些指标包括真阳性(TP)、假阳性(FP)、真阴性(TN)和假阴性(FN)等计数,并由此导出关键指标。敏感性(或召回率)衡量实际阳性中被正确识别的比例(TP/(TP+FN)),而特异性衡量实际阴性中被正确识别的比例(TN/(TN+FP))。精度(也称为阳性预测值(PPV))表示预测阳性中真正为阳性的比例(TP/(TP+FP))。相反,阴性预测值(NPV)表示预测阴性中真正为阴性的比例(TN/(TN+FN))。为平衡精度和召回率,F1分数作为这两个指标的调和平均数(2×(Precision×Recall)/(Precision+Recall))被用作反映整体性能的单一指标(211)。最后,有效的基准测试需要清晰地呈现结果,通常通过可视化和综合报告实现,以可解释的方式传达发现(214)。

        SV/CNV基准测试的关键挑战之一是,同一基因组改变可能被不同工具以不同方式报告,通常断点坐标、参考等位基因和事件描述各异,使跨方法的直接比较复杂化(211)。此外,比较参数的选择对基准测试结果有决定性影响。过于宽松的SV匹配阈值可能导致不同变异被过度合并、人为抬高性能指标、引入错误注释、将独特SV错误分类为共享、降低观察到的等位基因多样性并高估等位基因频率。相反,过于严格的阈值可能低估性能、遗漏有效注释,并降低下游分析的统计效力(215, 216)。CNV检测在重复基因组区域尤其具有挑战性,那里的变异通常由重复元件介导,使比对和变异解析复杂化(7)。尽管LR-seq的进展改善了对这些区域的特征描述,但在建立可靠真实集和比较变异检出方面仍存在重大挑战,主要原因是断点简并性(breakpoint degeneracy)以及存在多个同样有效的断点位置(29, 211, 217)。因此,许多基准数据集刻意排除高度重复区域(212, 218, 219)。尽管存在这些局限,真实集在多种情境下已非常有效,包括:使用精度和召回率等指标标准化跨工具性能评估;对结构复杂但医学相关区域的变异进行基准测试;以及促进基于组装的基准测试方法(如TT-Mars)的开发,后者可在基于坐标的比较不可靠的区域改进评估(211, 217)。最后,不同变异检出工具输出格式的异质性要求在进行比较分析前进行严格的标准化和归一化(211)。

        基准测试工具是基因组学中的必备框架,使研究人员能够使用标准化参考数据集和人工整理的真实集客观评估和比较SV与CNV检测方法的性能。Truvari被广泛认可用于SV检出集基准测试,通过将测试检出集与经验证的基准集进行严格比较,计算精度、召回率、F1分数和基因型一致性。它支持灵活的匹配标准,包括断点邻近性、大小相似性和序列相似性,使其适用于高分辨率和精度较低的SV检出(216)。SVanalyzer SVbenchmark采用先进的基于序列的匹配算法,能够对不同表示的SV检出进行稳健评估,尤其是在复杂或重复基因组区域,即使对于断点不精确或位于串联重复内的变异也能提高准确性(212)。TT-Mars引入了一种独特的基准测试范式,将SV检出与高质量的单倍型解析基因组组装进行比较,而非仅与常规真实集比较。TT-Mars不匹配预定义的检出集,而是评估每个SV检出所隐含的序列内容是否与组装单倍型中的实际序列一致,使其在缺乏人工整理真实集或由于基因组复杂性导致变异表示不明确的区域中尤其有价值(217)。

        CNVbenchmarkeR(220)及其增强版后继者CNVbenchmarkeR2(38)是基于R的综合框架,用于使用靶向下一代测序基因panel数据系统性地对CNV检出工具进行基准测试。CNVbenchmarkeR自动化执行和评估多种CNV检出器,包括DECoN、CoNVaDING、panelcn.MOPS、ExomeDepth和CODEX2,使用经多重连接依赖性探针扩增(MLPA)或基于阵列的比较基因组杂交(aCGH)金标准确认的验证数据集。它计算广泛的性能指标,包括敏感性、特异性、PPV、NPV和F1分数,并支持数据集特异性参数优化(220)。CNVbenchmarkeR2将这些能力扩展到对更广泛的CNV工具进行基准测试,如Atlas-CNV、ClinCNV、CNVkit和GATK-gCNV,并允许系统探索100多个工具参数、生成详细的图形报告以及评估整合多个工具结果的元检出器(meta-caller)策略(38)。两个框架都设计为用户友好且可适应,使研究和临床实验室能够直接使用其数据集对CNV检测工具进行基准测试。

        witty.er(https://github.com/Illumina/witty.er)是专门用于大片段变异的基准测试工具,明确设计用于评估SV和CNV检出器的性能。它由Illumina开发,将查询VCF与多种SV类型的真实VCF进行比较,应用灵活的位置和基因型匹配标准,并生成带注释VCF的详细基准测试统计。功能上,它类似于用于小变异基准测试的hap.py,是大片段变异检出性能评估的核心资源。

        svclassify作为基础性资源,支持创建高置信度基准SV检出集。它不是作为基准测试评估器,而是整合多平台测序证据并应用机器学习算法将候选SV分类为可能真阳性或假阳性。由此产生的高置信度SV集作为下游基准测试的真实集,确保可靠的性能评估并支持SV基准测试研究的可重复性(221)。SURVIVOR等支持工具在基准测试工作流程中发挥补充作用。SURVIVOR可模拟SV用于基准测试、合并和过滤SV检出集,并生成共识或模拟真实数据集用于评估,尽管它不生成标准化的基准测试报告(177)。

        通过提供指标计算、参数优化和跨工具比较的标准化工作流程,并与辅助工具整合用于检出集准备、模拟和验证,这些基准测试解决方案构成了CNV工具开发、方法选择以及研究和诊断环境中临床流程验证的基础要素。

CNV检出器独立基准测试研究

        生物信息学算法对于跨测序和微阵列平台的CNV检测至关重要。由于基准测试结果因数据类型、CNV大小和研究设计而差异显著,因此最好在平台背景下解读研究结果。鉴于可用算法众多,本综述综合了独立基准测试研究的发现,以提供不同情境下工具性能的无偏概述。详细的平台特异性基准测试结果和工具比较见附加文件3:注释S4(33-35, 37-39, 41, 43-45, 125, 163, 198, 220, 222-232)。

        总体而言,对跨多种技术的独立CNV检出器进行系统性基准测试揭示,没有任何单一工具或方法对所有实验情境、CNV大小或类型普遍最优。与其采用简单的重叠策略,对表现最佳的工具配对进行迭代合并可同时提高准确性和敏感性。然而,并非所有组合都有利,这凸显了进行系统性基准测试以识别协同组合的必要性。因此,实现稳健且准确的CNV检测需要多层次策略,包括仔细选择工具和调整参数,通常需要智能组合多个检出器并执行严格的后处理步骤。关键步骤包括优化参数(如窗口大小、探针选择、检出器-比对器配对和参考集组成),以及通过自定义过滤去除复发性假象并应用质控阈值。全面的后处理和质控检查,如覆盖深度分析、均一性评估和离群值检测,对于最小化假阳性和提高检出的CNV置信度至关重要(28, 35-39, 41, 43, 44, 198, 220, 229-232)。

        至关重要的是,仅靠计算预测不足以获得高置信度的CNV检出,尤其是对于罕见、小型且临床相关的变异。所有此类发现都应使用正交方法验证,从目视检查和二元比对图(BAM)审查到定量PCR(qPCR)、MLPA或基于阵列的方法,然后才能进行临床报告或解读。纳入检出后步骤,如基因分型细化、置信度评分(可能使用机器学习或贝叶斯模型)以及整合单倍型或替代参考信息,可进一步增强CNV检出的可靠性。最后,实际考量因素,如工作流可重复性(例如通过容器化和工作流管理器)、用户友好性、与现有临床或诊断流程的整合以及实验室生物信息学能力,对于成功的CNV分析同样重要(33, 34, 36-39, 43-45, 198, 220, 223, 225, 228, 229, 232)。虽然这些方法可有效检测CNV事件,但区分CNV发现与基因分型很重要。关联研究,尤其是涉及常见或多等位CNV的研究,需要跨队列准确估计拷贝数状态,正如基于队列和混合模型的方法(如CNest和GenomeSTRiP)所实现的(66, 233)。

3

质量控制与预处理

        QC是CNV检出前后的关键步骤,以确保下游分析前的样本和变异水平可靠性。在检出前阶段,基因分型率低、LRR或BAF指标离群、性别不一致或(对于测序数据)覆盖谱异常的个体被排除(21, 66, 135, 137-142, 234-238)。此阶段还应用额外校正,如GC含量校正以减轻杂交偏倚,以及主成分分析(PCA)以去除批次效应(20, 66, 86, 135-142, 234)。

        在检出后阶段(或等效地,对于基于信号的分析,在变异水平QC期间),去除低质量基因座。对于CNV检出,这包括排除CNV检出过多的样本(提示DNA质量差)、过滤噪声或多态性基因座的事件,以及移除未达到探针密度、读长深度或长度最低阈值的检出。对于基于信号的分析(如探针或外显子水平关联),包括排除覆盖度差、方差极端或可映射性低的单元。同一个体中相邻的CNV检出可合并以防止过度碎片化,而易于产生假象检出或体细胞变异的基因组问题区域,如着丝粒、端粒、免疫球蛋白基因座、T细胞受体区域、节段性重复和低拷贝重复,通常被排除(20, 21, 76, 86, 135-142, 234, 235, 239-241)。

        如前节所述,高置信度CNV通常通过多种检测方法的一致性、概率性质量评分或独立实验验证来定义(19, 20, 26, 76, 135, 239, 241)。此外,目视检查是验证CNV检出和过滤潜在假象的关键质控步骤。研究人员在此过程中通常检查几个关键特征,包括基于阵列数据中的信号强度模式和等位基因频率分布(如LRR和BAF),以及测序数据中的读长深度谱、分裂读长和不一致读长对信号以及局部比对质量,通常使用整合基因组学查看器(IGV)等基因组浏览器可视化。此外,使用UCSC基因组浏览器等资源评估基因组背景,例如评估与重复元件、节段性重复或低可映射性区域的邻近性,这可能提示潜在假象(20, 24, 26, 27, 86, 136, 140, 235, 241, 242)。

 

结论与未来展望

        CNV是塑造人类遗传多样性并影响多种疾病易感性的主要SV类别。测序技术和计算框架的实质性进展极大地增强了CNV检测;然而,在将CNV发现转化为可靠关联研究方面仍存在重要挑战。一个核心局限是没有任何单一CNV检出工具能在所有实验环境中持续优于其他工具;因此,多种算法的联合使用仍是CNV检测最稳健的策略。除检测外,严格的QC、审慎的CNV分组以及适当的关联检验和效应建模统计框架对于确保有效且可解读的结果至关重要。综合而言,这些考量凸显了CNV分析的复杂性,并强调需要持续的方法学创新以充分阐明CNV对健康和疾病的贡献。

        CNV分析领域预计将经历深刻变革,由参考基因组组装、整合方法、基准测试标准和计算工具的进展驱动。综合性参考基因组的发展,尤其是端粒到端粒(T2T)CHM13组装,提供了无间隙的人类基因组序列,纠正了结构错误并增加了约200 Mb此前未解析的序列,其中许多位于着丝粒、近端着丝粒和端粒区域,并包含额外的蛋白编码基因(301)。在此基础上,通过基因组图谱(genome graphs)代表多样化人群的人类参考泛基因组(pangenome)的建立,为更好地表征重复区域和复杂SV提供了有力手段(302-305)。人类泛基因组参考联盟(Human Pangenome Reference Consortium)已产出包含大量常染色质序列和GRCh38中缺失SV的分阶段二倍体组装草案(306)。GraphTyper(307, 308)等工具利用泛基因组图谱进行人群规模SV基因分型,体现了将基因组图谱表示整合到大规模研究中如何克服参考偏倚、增强读长比对并促进跨多样化人群的准确SV基因分型。然而,将现有注释和基准迁移到基于图谱的参考面临挑战,尤其是在将基因定义和调控注释适配到这一新框架方面(2)。

        与此同时,LR-seq已成为SV发现和CNV分析的变革性技术。Beyter等(4)的一项人群规模LR-seq研究识别出每个个体超过22,000个SV,是短读长测序发现数量的三到五倍,凸显了敏感性方面的显著提升,尤其是对于串联重复相关SV。此外,将LR-seq衍生的SV与填补(imputation)整合进大型队列已实现功能性预测,如识别出与LDL胆固醇水平降低相关的PCSK9罕见缺失以及与身高相关的ACAN多等位重复(4)。最近,Bai等(309)利用482个单倍型解析长读长组装构建了参考panel。他们开发了在线填补工具ImputeSV,旨在从SNP数据填补SV和串联重复变异。该工具被应用于在英国生物样本库456,643名参与者中填补54,578个常见SV。该研究识别出2,624个性状中的17,335个SV-性状关联,并估计SV至少贡献了与复杂性状相关的常见遗传变异的4.7%。这些发现强调了LR-seq作为泛基因组方法关键补充的重要性,提供了更好捕捉人群多样性和疾病相关变异的综合性SV目录。

        未来方向还强调在统一分析框架内整合CNV和SNP关联研究。许多CNV难以被SNP标记或出现在不同单倍型上,凸显了CNV特异性分析的必要性。CNest等方法已促进基于与SNP信号重叠情况对CNV关联进行分类,区分CNV独有、CNV等位基因、SNP-CNV邻近和SNP-CNV远距离关联。SNP和CNV关联的联合建模预计将揭示共享遗传机制、增强因果基因的优先级排序,并细化多基因风险评分(PGS)预测,前提是建立SNP与CNV之间的LD图谱(27, 66, 310)。

        未来的CNV研究将越来越多地采用多模态整合策略,将CNV数据与其他分子层面(如转录组、表观基因组和染色质拓扑)相结合,以更好地理解SV/CNV的功能影响。近期研究强调了罕见胚系SV如何破坏高表达和突变受限的基因、改变拓扑关联结构域(TADs)等调控域,并在组织特异性情境中使基因表达失调(3, 28)。例如,将CNV数据与来自疾病相关组织的RNA测序和表观遗传谱整合已证明,单例(singleton)基因破坏性胚系CNV优先影响起源组织表达的基因,可能导致儿科肿瘤中可观察到的下游转录失调。此外,与组织特异性染色质边界重叠的非编码CNV与三维基因组架构破坏相关,将CNV与调控网络改变联系起来(311)。

        CNV检测的基准测试和评估仍然具有挑战性,尤其是在重复基因组区域。新兴工具(如将SV检出与单倍型解析组装进行比较的TT-Mars)正在帮助应对重复区域的基准测试挑战(217)。此外,nf-core/variantbenchmarking流程通过整合多样化真实集并支持跨组装比较(包括CNV特异性基准)来简化评估(312)。

        除基准测试外,迫切需要标准化报告指南和数据定义,以确保可重复性和互操作性,最终促进将CNV发现纳入GWAS目录等公共资源。在此背景下,因果推断方法(如转录组-wide孟德尔随机化)尤其有价值,因为它们为将CNV与功能性基因表达改变和表型效应联系起来提供了稳健框架,从而加强CNV关联结果正式整合到目录和下游分析(如PGS和药物靶点优先级排序)所需的证据(27, 82, 313)。

        创新性工作流和单倍型感知方法已变革大规模CNV发现与分析。例如,CNest工作流可在非常大的队列中从NGS读长深度实现高分辨率CNV检测,并已在英国生物样本库中发现数百个新关联(66)。至关重要的是,CNest遵循全球基因组学与健康联盟(GA4GH)标准,确保人群规模研究中的互操作性和标准化数据交换(314)。与此互补,HI-CNV等单倍型知情方法通过利用生物样本库规模数据集中的共享单倍型大幅提高检测敏感性,识别的每个体CNV数量是早期方法的六倍以上(20)。最后,SHAPEIT5等统计定相工具的进展提供了CNV分析所需的单倍型支架,使得能够探索相位依赖的等位基因系列及其下游表型后果(315)。

 

结论

        本研究呈现了中国人群遗传性心肌病和心律失常遗传检测的诊断阳性率为25.4%。应积极考虑对老年人群以及一级亲属进行遗传检测。鉴于大多数基因检测结果可指导治疗和预后评估,复杂结果应进行进一步的致病证据分析以得出可靠结论。这一方法可为临床决策提供更多见解。

参考文献:

[1] Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, Korbel JO, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68-74.

[2] Collins RL, Talkowski ME. Diversity and consequences of structural variation in the human genome. Nat Rev Genet. 2025.

[3] Spielmann M, Lupiáñez DG, Mundlos S. Structural variation in the 3D genome. Nat Rev Genet. 2018;19(7):453-67.

[4] Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H, et al. Long-read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits. Nature Genetics. 2021;53(6):779-86.

[5] Zarrei M, MacDonald J, Merico D, Scherer S. A copy number variation map of the human genome. Nature Reviews Genetics. 2015;16:172-83.

[6] Abel HJ, Larson DE, Regier AA, Chiang C, Das I, Kanchi KL, et al. Mapping and characterization of structural variation in 17,795 human genomes. Nature. 2020;583(7814):83-9.

[7] Carvalho CM, Lupski JR. Mechanisms underlying structural variant formation in genomic disorders. Nat Rev Genet. 2016;17(4):224-38.

[8] Harris L, McDonagh EM, Zhang X, Fawcett K, Foreman A, Daneck P, et al. Genome-wide association testing beyond SNPs. Nat Rev Genet. 2025;26(3):156-70.

[9] Hurles ME, Dermitzakis ET, Tyler-Smith C. The functional impact of structural variation in humans. Trends Genet. 2008;24(5):238-45.

[10] Jakubosky D, D’Antonio M, Bonder MJ, Smail C, Donovan MKR, Young Greenwald WW, et al. Properties of structural variants and short tandem repeats associated with gene expression and complex traits. Nature Communications. 2020;11(1):2927. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[11] Scott AJ, Chiang C, Hall IM. Structural variants are a major source of gene expression differences in humans and often affect multiple nearby genes. Genome Res. 2021;31(12):2249-57.

[12] Gunning AC, Strucinska K, Muñoz Oreja M, Parrish A, Caswell R, Stals KL, et al. Recurrent De Novo NAHR Reciprocal Duplications in the ATAD3 Gene Cluster Cause a Neurogenetic Trait with Perturbed Cholesterol and Mitochondrial Metabolism. Am J Hum Genet. 2020;106(2):272-9.

[13] Harel T, Lupski JR. Genomic disorders 20 years on-mechanisms for clinical manifestations. Clin Genet. 2018;93(3):439-49.

[14] Collins RL, Glessner JT, Porcu E, Lepamets M, Brandon R, Lauricella C, et al. A cross-disorder dosage sensitivity map of the human genome. Cell. 2022;185(16):3041- 55.e25.

[15] Smolen C, Girirajan S. The gene dose makes the disease. Cell. 2022;185(16):2850-2.

[16] Girirajan S, Rosenfeld JA, Coe BP, Parikh S, Friedman N, Goldstein A, et al. Phenotypic heterogeneity of genomic disorders and rare copy-number variants. N Engl J Med. 2012;367(14):1321-31.

[17] Albers CA, Paul DS, Schulze H, Freson K, Stephens JC, Smethurst PA, et al. Compound inheritance of a low-frequency regulatory SNP and a rare null mutation in exon-junction complex subunit RBM8A causes TAR syndrome. Nat Genet. 2012;44(4):435-9, s1-2.

[18] Aguirre M, Rivas MA, Priest J. Phenome-wide Burden of Copy-Number Variation in the UK Biobank. Am J Hum Genet. 2019;105(2):373-83.

[19] Macé A, Tuke MA, Deelen P, Kristiansson K, Mattsson H, Nõukas M, et al. CNV- association meta-analysis in 191,161 European adults reveals new loci associated with anthropometric traits. Nat Commun. 2017;8(1):744.

[20] Hujoel MLA, Sherman MA, Barton AR, Mukamel RE, Sankaran VG, Terao C, et al. Influences of rare copy-number variation on human complex traits. Cell. 2022;185(22):4233-48.e27.

[21] Marshall CR, Howrigan DP, Merico D, Thiruvahindrapuram B, Wu W, Greer DS, et al. Contribution of copy number variants to schizophrenia from a genome-wide study of 41,321 subjects. Nat Genet. 2017;49(1):27-35.

[22] Trost B, Thiruvahindrapuram B, Chan AJS, Engchuan W, Higginbotham EJ, Howe JL, et al. Genomic architecture of autism from comprehensive whole-genome sequence annotation. Cell. 2022;185(23):4409-27.e18.

[23] Gu Z, Gu L, Eils R, Schlesner M, Brors B. circlize implements and enhances circular visualization in R. Bioinformatics. 2014;30(19):2811-2.

[24] Fawcett KA, Demidov G, Shrine N, Paynton ML, Ossowski S, Sayers I, et al. Exome-wide analysis of copy number variation shows association of the human leukocyte antigen region with asthma in UK Biobank. BMC Med Genomics. 2022;15(1):119.

[25] Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562(7726):203-9. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[26] Auwerx C, Jõeloo M, Sadler MC, Tesio N, Ojavee S, Clark CJ, et al. Rare copy- number variants as modulators of common disease susceptibility. Genome Med. 2024;16(1):5.

[27] Auwerx C, Lepamets M, Sadler MC, Patxot M, Stojanov M, Baud D, et al. The individual and global impact of copy-number variants on complex human traits. Am J Hum Genet. 2022;109(4):647-68.

[28] Ho SS, Urban AE, Mills RE. Structural variation in the sequencing era. Nat Rev Genet. 2020;21(3):171-89.

[29] Mahmoud M, Gobet N, Cruz-Dávalos DI, Mounier N, Dessimoz C, Sedlazeck FJ. Structural variant calling: the long and the short of it. Genome Biol. 2019;20(1):246.

[30] Audano PA, Sulovari A, Graves-Lindsay TA, Cantsilieris S, Sorensen M, Welch AE, et al. Characterizing the Major Structural Variant Alleles of the Human Genome. Cell. 2019;176(3):663-75.e19.

[31] Ebert P, Audano PA, Zhu Q, Rodriguez-Martin B, Porubsky D, Bonder MJ, et al. Haplotype-resolved diverse human genomes and integrated analysis of structural variation. Science. 2021;372(6537).

[32] Logsdon GA, Ebert P, Audano PA, Loftus M, Porubsky D, Ebler J, et al. Complex genetic variation in nearly complete human genomes. Nature. 2025;644(8076):430-41.

[33] Gabrielaite M, Torp MH, Rasmussen MS, Andreu-Sánchez S, Vieira FG, Pedersen CB, et al. A Comparison of Tools for Copy-Number Variation Detection in Germline Whole Exome and Whole Genome Sequencing Data. Cancers (Basel). 2021;13(24).

[34] van Baardwijk MN, Heijnen L, Zhao H, Baudis M, Stubbs AP. A systematic benchmark of copy number variation detection tools for high density SNP genotyping arrays. Genomics. 2024;116(6):110962.

[35] Whitford W, Lehnert K, Snell RG, Jacobsen JC. Evaluation of the performance of copy number variant prediction tools for the detection of deletions from whole genome sequencing data. J Biomed Inform. 2019;94:103174.

[36] Yao R, Zhang C, Yu T, Li N, Hu X, Wang X, et al. Evaluation of three read-depth based CNV detection tools using whole-exome sequencing data. Mol Cytogenet. 2017;10:30.

[37] De La Vega FM, Irvine SA, Anur P, Potts K, Kraft L, Torres R, et al. Benchmarking of germline copy number variant callers from whole genome sequencing data for clinical applications. Bioinform Adv. 2025;5(1):vbaf071.

[38] Munté E, Roca C, Del Valle J, Feliubadaló L, Pineda M, Gel B, et al. Detection of germline CNVs from gene panel data: benchmarking the state of the art. Brief Bioinform. 2024;26(1).

[39] Yuan N, Jia P. Comprehensive assessment of long-read sequencing platforms and calling algorithms for detection of copy number variation. Brief Bioinform. 2024;25(5).

[40] Song M, Ma S, Wang G, Wang Y, Yang Z, Xie B, et al. Benchmarking copy number aberrations inference tools using single-cell multi-omics datasets. Brief Bioinform. 2025;26(2).

[41] Kosugi S, Momozawa Y, Liu X, Terao C, Kubo M, Kamatani Y. Comprehensive evaluation of structural variation detection algorithms for whole genome sequencing. Genome Biol. 2019;20(1):117. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[42] Majidian S, Agustinho DP, Chin CS, Sedlazeck FJ, Mahmoud M. Genomic variant benchmark: if you cannot measure it, you cannot improve it. Genome Biol. 2023;24(1):221.

[43] Coutelier M, Holtgrewe M, Jäger M, Flöttman R, Mensah MA, Spielmann M, et al. Combining callers improves the detection of copy number variants from whole- genome sequencing. Eur J Hum Genet. 2022;30(2):178-86.

[44] Trost B, Walker S, Wang Z, Thiruvahindrapuram B, MacDonald JR, Sung WWL, et al. A Comprehensive Workflow for Read Depth-Based Identification of Copy-Number Variation from Whole-Genome Sequence Data. Am J Hum Genet. 2018;102(1):142-55.

[45] Gordeeva V, Sharova E, Babalyan K, Sultanov R, Govorun VM, Arapidi G. Benchmarking germline CNV calling tools from exome sequencing data. Scientific Reports. 2021;11(1):14416.

[46] Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015;4:7.

[47] Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559-75.

[48] Mbatchou J, Barnard L, Backman J, Marcketta A, Kosmicki JA, Ziyatdinov A, et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat Genet. 2021;53(7):1097-103.

[49] Glessner JT, Li J, Liu Y, Khan M, Chang X, Sleiman PMA, et al. ParseCNV2: efficient sequencing tool for copy number variation genome-wide association studies. European Journal of Human Genetics. 2023;31(3):304-12.

[50] Kim JH, Hu HJ, Yim SH, Bae JS, Kim SY, Chung YJ. CNVRuler: a copy number variation-based case-control association analysis tool. Bioinformatics. 2012;28(13):1790-2.

[51] Forer L, Schönherr S, Weissensteiner H, Haider F, Kluckner T, Gieger C, et al. CONAN: copy number variation analysis software for genome-wide association studies. BMC Bioinformatics. 2010;11(1):318.

[52] Glessner JT, Li J, Hakonarson H. ParseCNV integrative copy number variation association software with quality tracking. Nucleic Acids Res. 2013;41(5):e64.

[53] Chen X, Fang LT, Chen Z, Chen W, Wu H, Zhu B, et al. A benchmarking study of copy number variation inference methods using single-cell RNA-sequencing data. Precision Clinical Medicine. 2025.

[54] Mallory XF, Edrisi M, Navin N, Nakhleh L. Assessing the performance of methods for copy number aberration detection from single-cell DNA sequencing data. PLoS Comput Biol. 2020;16(7):e1008012.

[55] Schmid KT, Symeonidi A, Hlushchenko D, Richter ML, Colomé-Tatché M. Benchmarking scRNA-seq copy number variation callers. bioRxiv. 2024:2024.12.18.629083.

[56] De Falco A, Caruso F, Su X-D, Iavarone A, Ceccarelli M. A variational algorithm to detect the clonal copy number substructure of tumors from scRNA-seq data. Nature Communications. 2023;14(1):1074. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[57] Fan J, Lee HO, Lee S, Ryu DE, Lee S, Xue C, et al. Linking transcriptional and genetic tumor heterogeneity through allele analysis of single-cell RNA-seq data. Genome Res. 2018;28(8):1217-27.

[58] Gao R, Bai S, Henderson YC, Lin Y, Schalck A, Yan Y, et al. Delineating copy number and clonal substructure in human tumors from single-cell transcriptomes. Nat Biotechnol. 2021;39(5):599-608.

[59] Gao T, Soldatov R, Sarkar H, Kurkiewicz A, Biederstedt E, Loh PR, et al. Haplotype-aware analysis of somatic copy number variations from single-cell transcriptomes. Nat Biotechnol. 2023;41(3):417-26.

[60] Harmanci AS, Harmanci A, Zhou X. CaSpER identifies and visualizes CNV events by integrative analysis of single-cell or bulk RNA-sequencing data. Nature Communications. 2020;11.

[61] Müller S, Cho A, Liu SJ, Lim DA, Diaz A. CONICS integrates scRNA-seq with DNA sequencing to map gene expression to tumor sub-clones. Bioinformatics. 2018;34(18):3217-9.

[62] Shao X, Lv N, Liao J, Long J, Xue R, Ai N, et al. Copy number variation is highly correlated with differential gene expression: a pan-cancer study. BMC Med Genet. 2019;20(1):175.

[63] Yi K, Ju YS. Patterns and mechanisms of structural variations in human cancer. Experimental & Molecular Medicine. 2018;50(8):1-11.

[64] Maury EA, Sherman MA, Genovese G, Gilgenast TG, Kamath T, Burris SJ, et al. Schizophrenia-associated somatic copy-number variants from 12,834 cases reveal recurrent NRXN1 and ABCB11 disruptions. Cell Genom. 2023;3(8):100356.

[65] McConnell MJ, Lindberg MR, Brennand KJ, Piper JC, Voet T, Cowing-Zitron C, et al. Mosaic copy number variation in human neurons. Science. 2013;342(6158):632-7.

[66] Fitzgerald T, Birney E. CNest: A novel copy number association discovery method uncovers 862 new associations from 200,629 whole-exome sequence datasets in the UK Biobank. Cell Genomics. 2022;2(8):100167.

[67] Pinto D, Darvishi K, Shi X, Rajan D, Rigler D, Fitzgerald T, et al. Comprehensive assessment of array-based platforms and calling algorithms for detection of copy number variants. Nat Biotechnol. 2011;29(6):512-20.

[68] Wiszniewska J, Bi W, Shaw C, Stankiewicz P, Kang S-HL, Pursley AN, et al. Combined array CGH plus SNP genome analyses in a single assay for optimized clinical testing. European Journal of Human Genetics. 2014;22(1):79-87.

[69] Verlouw JAM, Clemens E, de Vries JH, Zolk O, Verkerk AJMH, am Zehnhoff- Dinnesen A, et al. A comparison of genotyping arrays. European Journal of Human Genetics. 2021;29(11):1611-24.

[70] Li YR, Glessner JT, Coe BP, Li J, Mohebnasab M, Chang X, et al. Rare copy number variants in over 100,000 European ancestry subjects reveal multiple disease associations. Nat Commun. 2020;11(1):255.

[71] Korn JM, Kuruvilla FG, McCarroll SA, Wysoker A, Nemesh J, Cawley S, et al. Integrated genotype calling and association analysis of SNPs, common copy number polymorphisms and rare CNVs. Nat Genet. 2008;40(10):1253-60. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[72] Zhang F, Gu W, Hurles ME, Lupski JR. Copy Number Variation in Human Health, Disease, and Evolution. Annual Review of Genomics and Human Genetics. 2009;10(Volume 10, 2009):451-81.

[73] Fu W, Zhang F, Wang Y, Gu X, Jin L. Identification of Copy Number Variation Hotspots in Human Populations. The American Journal of Human Genetics. 2010;87(4):494-504.

[74] Brandler William M, Antaki D, Gujral M, Noor A, Rosanio G, Chapman Timothy R, et al. Frequency and Complexity of De Novo Structural Mutation in Autism. The American Journal of Human Genetics. 2016;98(4):667-79.

[75] Belyeu JR, Brand H, Wang H, Zhao X, Pedersen BS, Feusier J, et al. De novo structural mutation rates and gamete-of-origin biases revealed through genome sequencing of 2,396 families. The American Journal of Human Genetics. 2021;108(4):597-607.

[76] Null M, Yilmaz F, Astling D, Yu HC, Cole JB, Hallgrímsson B, et al. Genome-wide analysis of copy number variants and normal facial variation in a large cohort of Bantu Africans. HGG Adv. 2022;3(1):100082.

[77] Craddock N, Hurles ME, Cardin N, Pearson RD, Plagnol V, Robson S, et al. Genome-wide association study of CNVs in 16,000 cases of eight common diseases and 3,000 shared controls. Nature. 2010;464(7289):713-20.

[78] Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, et al. A structural variation reference for medical and population genetics. Nature. 2020;581(7809):444-51.

[79] Campbell CD, Eichler EE. Properties and rates of germline mutations in humans. Trends Genet. 2013;29(10):575-84.

[80] Handsaker RE, Van Doren V, Berman JR, Genovese G, Kashin S, Boettger LM, et al. Large multiallelic copy number variations in humans. Nat Genet. 2015;47(3):296-

[81] Jakubosky D, Smith EN, D’Antonio M, Jan Bonder M, Young Greenwald WW, D’Antonio-Chronowska A, et al. Discovery and quality analysis of a comprehensive set of structural variants and short tandem repeats. Nat Commun. 2020;11(1):2928.

[82] Zamariolli M, Auwerx C, Sadler MC, van der Graaf A, Lepik K, Schoeler T, et al. The impact of 22q11.2 copy-number variants on human traits in the general population. Am J Hum Genet. 2023;110(2):300-13.

[83] Hanssen R, Auwerx C, Jõeloo M, Sadler MC, Henning E, Keogh J, et al. Chromosomal deletions on 16p11.2 encompassing SH2B1 are associated with accelerated metabolic disease. Cell Rep Med. 2023;4(8):101155.

[84] Cooper DN, Krawczak M, Polychronakos C, Tyler-Smith C, Kehrer-Sawatzki H. Where genotype is not predictive of phenotype: towards an understanding of the molecular basis of reduced penetrance in human inherited disease. Human Genetics. 2013;132(10):1077-130.

[85] Crawford K, Bracher-Smith M, Owen D, Kendall KM, Rees E, Pardiñas AF, et al. Medical consequences of pathogenic CNVs in adults: analysis of the UK Biobank. J Med Genet. 2019;56(3):131-8. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[86] Rodríguez-López J, Flórez G, Blanco V, Pereiro C, Fernández JM, Fariñas E, et al. Genome wide analysis of rare copy number variations in alcohol abuse or dependence. J Psychiatr Res. 2018;103:212-8.

[87] Barnes C, Plagnol V, Fitzgerald T, Redon R, Marchini J, Clayton D, et al. A robust statistical method for case-control association testing with copy number variation. Nat Genet. 2008;40(10):1245-52.

[88] da Silva V, Ramos M, Groenen M, Crooijmans R, Johansson A, Regitano L, et al. CNVRanger: association analysis of CNVs with gene expression and quantitative phenotypes. Bioinformatics. 2020;36(3):972-3.

[89] Comprehensive gene panels provide advantages over clinical exome sequencing for Mendelian diseases. Genome Biol. 2015;16(1):134.

[90] Ashley EA. Towards precision medicine. Nat Rev Genet. 2016;17(9):507-22.

[91] Chaisson MJ, Huddleston J, Dennis MY, Sudmant PH, Malig M, Hormozdiari F, et al. Resolving the complexity of the human genome using single-molecule sequencing. Nature. 2015;517(7536):608-11.

[92] De Falco A, Caruso F, Su XD, Iavarone A, Ceccarelli M. A variational algorithm to detect the clonal copy number substructure of tumors from scRNA-seq data. Nat Commun. 2023;14(1):1074.

[93] Goldfeder RL, Priest JR, Zook JM, Grove ME, Waggott D, Wheeler MT, et al. Medical implications of technical accuracy in genome sequencing. Genome Med. 2016;8(1):24.

[94] Goodwin S, McPherson JD, McCombie WR. Coming of age: ten years of next- generation sequencing technologies. Nat Rev Genet. 2016;17(6):333-51.

[95] Gordeeva V, Sharova E, Arapidi G. Progress in Methods for Copy Number Variation Profiling. Int J Mol Sci. 2022;23(4).

[96] Huddleston J, Chaisson MJP, Steinberg KM, Warren W, Hoekzema K, Gordon D, et al. Discovery and genotyping of structural variation from long-read haploid genome sequence data. Genome Res. 2017;27(5):677-85.

[97] Iafrate AJ, Feuk L, Rivera MN, Listewnik ML, Donahoe PK, Qi Y, et al. Detection of large-scale variation in the human genome. Nat Genet. 2004;36(9):949-51.

[98] Klein CJ, Foroud TM. Neurology Individualized Medicine: When to Use Next- Generation Sequencing Panels. Mayo Clin Proc. 2017;92(2):292-305.

[99] Merker JD, Wenger AM, Sneddon T, Grove M, Zappala Z, Fresard L, et al. Long- read genome sequencing identifies causal structural variation in a Mendelian disease. Genet Med. 2018;20(1):159-63.

[100] Miller DT, Adam MP, Aradhya S, Biesecker LG, Brothman AR, Carter NP, et al. Consensus statement: chromosomal microarray is a first-tier clinical diagnostic test for individuals with developmental disabilities or congenital anomalies. Am J Hum Genet. 2010;86(5):749-64.

[101] Schaefer GB, Mendelsohn NJ. Clinical genetics evaluation in identifying the etiology of autism spectrum disorders: 2013 guideline revisions. Genet Med. 2013;15(5):399-407.

[102] Sebat J, Lakshmi B, Troge J, Alexander J, Young J, Lundin P, et al. Large-scale copy number polymorphism in the human genome. Science. 2004;305(5683):525-8. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[103] Song JHT, Lowe CB, Kingsley DM. Characterization of a Human-Specific Tandem Repeat Associated with Bipolar Disorder and Schizophrenia. Am J Hum Genet. 2018;103(3):421-30.

[104] Zhao M, Wang Q, Wang Q, Jia P, Zhao Z. Computational tools for copy number variation (CNV) detection using next-generation sequencing data: features and perspectives. BMC Bioinformatics. 2013;14(11):S1.

[105] Ahsan MU, Liu Q, Perdomo JE, Fang L, Wang K. A survey of algorithms for the detection of genomic structural variants from long-read sequencing data. Nat Methods. 2023;20(8):1143-58.

[106] Colella S, Yau C, Taylor JM, Mirza G, Butler H, Clouston P, et al. QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data. Nucleic Acids Res. 2007;35(6):2013-25.

[107] Cretu Stancu M, van Roosmalen MJ, Renkens I, Nieboer MM, Middelkamp S, de Ligt J, et al. Mapping and phasing of structural variation in patient genomes using nanopore sequencing. Nature Communications. 2017;8(1):1326.

[108] Darvishi K. Application of Nexus copy number software for CNV detection and analysis. Curr Protoc Hum Genet. 2010;Chapter 4:Unit 4.14.1-28.

[109] Ding H, Luo J. MAMnet: detecting and genotyping deletions and insertions based on long reads and a deep learning approach. Brief Bioinform. 2022;23(5).

[110] English AC, Salerno WJ, Reid JG. PBHoney: identifying genomic variants via long- read discordance and interrupted mapping. BMC Bioinformatics. 2014;15(1):180.

[111] Goel M, Sun H, Jiao W-B, Schneeberger K. SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biology. 2019;20(1):277.

[112] Gong L, Wong CH, Cheng WC, Tjong H, Menghi F, Ngan CY, et al. Picky comprehensively detects high-resolution structural variants in nanopore long reads. Nat Methods. 2018;15(6):455-60.

[113] Heller D, Vingron M. SVIM-asm: structural variant detection from haploid and diploid genome assemblies. Bioinformatics. 2020;36(22-23):5519-21.

[114] Jeng XJ, Cai TT, Li H. Optimal Sparse Segment Identification with Application in Copy Number Variation Analysis. J Am Stat Assoc. 2010;105(491):1156-66.

[115] Jiang T, Liu Y, Jiang Y, Li J, Gao Y, Cui Z, et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biology. 2020;21(1):189.

[116] Killick R, Fearnhead P, Eckley IA. Optimal detection of changepoints with a linear computational cost. Journal of the American Statistical Association. 2012;107(500):1590-8.

[117] Lin J, Wang S, Audano PA, Meng D, Flores JI, Kosters W, et al. SVision: a deep learning approach to resolve complex structural variants. Nat Methods. 2022;19(10):1230-3.

[118] Luo J, Ding H, Shen J, Zhai H, Wu Z, Yan C, et al. BreakNet: detecting deletions using long reads and a deep learning approach. BMC Bioinformatics. 2021;22(1):577.

[119] Magi A, Tattini L, Pippucci T, Torricelli F, Benelli M. Read count approach for DNA copy number variants detection. Bioinformatics. 2012;28(4):470-8.

[120] Niu YS, Zhang H. THE SCREENING AND RANKING ALGORITHM TO DETECT DNA COPY NUMBER VARIATIONS. Ann Appl Stat. 2012;6(3):1306-26. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[121] Olshen AB, Venkatraman ES, Lucito R, Wigler M. Circular binary segmentation for the analysis of array-based DNA copy number data. Biostatistics. 2004;5(4):557-72.

[122] Pinto D, Pagnamenta AT, Klei L, Anney R, Merico D, Regan R, et al. Functional impact of global rare copy number variation in autism spectrum disorders. Nature. 2010;466(7304):368-72.

[123] Pique-Regi R, Cáceres A, González JR. R-Gada: a fast and flexible pipeline for copy number analysis in association studies. BMC Bioinformatics. 2010;11:380.

[124] Pirooznia M, Goes FS, Zandi PP. Whole-genome CNV analysis: advances in computational approaches. Frontiers in Genetics. 2015;6.

[125] Roy S, Motsinger Reif A. Evaluation of calling algorithms for array-CGH. Front Genet. 2013;4:217.

[126] Sedlazeck FJ, Rescheneder P, Smolka M, Fang H, Nattestad M, von Haeseler A, et al. Accurate detection of complex structural variations using single-molecule sequencing. Nature Methods. 2018;15(6):461-8.

[127] Smolka M, Paulin LF, Grochowski CM, Horner DW, Mahmoud M, Behera S, et al. Detection of mosaic and population-level structural variants with Sniffles2. Nature Biotechnology. 2024;42(10):1571-80.

[128] Tham CY, Tirado-Magallanes R, Goh Y, Fullwood MJ, Koh BTH, Wang W, et al. NanoVar: accurate characterization of patients’ genomic structural variants using low- depth nanopore sequencing. Genome Biology. 2020;21(1):56.

[129] Tibshirani R, Wang P. Spatial smoothing and hot spot detection for CGH data using the fused lasso. Biostatistics. 2008;9(1):18-29.

[130] Wang K, Li M, Hadley D, Liu R, Glessner J, Grant SF, et al. PennCNV: an integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data. Genome Res. 2007;17(11):1665-74.

[131] Wenger AM, Peluso P, Rowell WJ, Chang PC, Hall RJ, Concepcion GT, et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol. 2019;37(10):1155-62.

[132] Zhang Z, Cheng H, Hong X, Di Narzo AF, Franzen O, Peng S, et al. EnsembleCNV: an ensemble machine learning algorithm to identify and genotype copy number variation using SNP array data. Nucleic Acids Res. 2019;47(7):e39.

[133] Zhang Z, Lange K, Ophoff R, Sabatti C. RECONSTRUCTING DNA COPY NUMBER BY PENALIZED ESTIMATION AND IMPUTATION. Ann Appl Stat. 2010;4(4):1749-73.

[134] Gel B, Serra E. karyoploteR: an R/Bioconductor package to plot customizable genomes displaying arbitrary data. Bioinformatics. 2017;33(19):3088-90.

[135] Lo Faro V, Ten Brink JB, Snieder H, Jansonius NM, Bergen AA. Genome-wide CNV investigation suggests a role for cadherin, Wnt, and p53 pathways in primary open- angle glaucoma. BMC Genomics. 2021;22(1):590.

[136] Hakkaart C, Pearson JF, Marquart L, Dennis J, Wiggins GAR, Barnes DR, et al. Copy number variants as modifiers of breast cancer risk for BRCA1/BRCA2 pathogenic variant carriers. Communications Biology. 2022;5(1):1061.

[137] Dennis J, Tyrer JP, Walker LC, Michailidou K, Dorling L, Bolla MK, et al. Rare germline copy number variants (CNVs) and breast cancer risk. Commun Biol. 2022;5(1):65. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[138] Kendall KM, Rees E, Escott-Price V, Einon M, Thomas R, Hewitt J, et al. Cognitive Performance Among Carriers of Pathogenic Copy Number Variants: Analysis of 152,000 UK Biobank Subjects. Biol Psychiatry. 2017;82(2):103-10.

[139] Kikuchi M, Kobayashi K, Nishida N, Sawai H, Sugiyama M, Mizokami M, et al. Genome-wide copy number variation analysis of hepatitis B infection in a Japanese population. Hum Genome Var. 2021;8(1):22.

[140] Li Z, Chen J, Xu Y, Yi Q, Ji W, Wang P, et al. Genome-wide Analysis of the Role of Copy Number Variation in Schizophrenia Risk in Chinese. Biol Psychiatry. 2016;80(4):331-7.

[141] Tansey KE, Rees E, Linden DE, Ripke S, Chambert KD, Moran JL, et al. Common alleles contribute to schizophrenia in CNV carriers. Mol Psychiatry. 2016;21(8):1085-9.

[142] Yuan J, Hu J, Li Z, Zhang F, Zhou D, Jin C. A replication study of schizophrenia- related rare copy number variations in a Han Southern Chinese population. Hereditas. 2017;154(1):2.

[143] Abyzov A, Urban AE, Snyder M, Gerstein M. CNVnator: an approach to discover, genotype, and characterize typical and atypical CNVs from family and population genome sequencing. Genome Res. 2011;21(6):974-84.

[144] Anjum S, Morganella S, D’Angelo F, Iavarone A, Ceccarelli M. VEGAWES: variational segmentation on whole exome sequencing for copy number detection. BMC Bioinformatics. 2015;16:315.

[145] Cabello-Aguilar S, Vendrell JA, Van Goethem C, Brousse M, Gozé C, Frantz L, et al. ifCNV: A novel isolation-forest-based package to detect copy-number variations from various targeted NGS datasets. Mol Ther Nucleic Acids. 2022;30:174-83.

[146] Fromer M, Moran JL, Chambert K, Banks E, Bergen SE, Ruderfer DM, et al. Discovery and statistical genotyping of copy-number variation from whole-exome sequencing depth. Am J Hum Genet. 2012;91(4):597-607.

[147] Magi A, Benelli M, Marseglia G, Nannetti G, Scordo MR, Torricelli F. A shifting level model algorithm that identifies aberrations in array-CGH data. Biostatistics. 2010;11(2):265-80.

[148] Magi A, Benelli M, Yoon S, Roviello F, Torricelli F. Detecting common copy number variants in high-throughput sequencing data by using JointSLM algorithm. Nucleic Acids Res. 2011;39(10):e65.

[149] Magi A, Pippucci T, Sidore C. XCAVATOR: accurate detection and genotyping of copy number variants from second and third generation whole-genome sequencing experiments. BMC Genomics. 2017;18(1):747.

[150] Magi A, Tattini L, Cifola I, D’Aurizio R, Benelli M, Mangano E, et al. EXCAVATOR: detecting copy number variants from whole-exome sequencing data. Genome Biol. 2013;14(10):R120.

[151] Miller CA, Hampton O, Coarfa C, Milosavljevic A. ReadDepth: a parallel R package for detecting copy number alterations from short sequencing reads. PLoS One. 2011;6(1):e16327.

[152] Morganella S, Cerulo L, Viglietto G, Ceccarelli M. VEGA: variational segmentation for copy number detection. Bioinformatics. 2010;26(24):3020-7. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[153] Mumford D, Shah J. Optimal approximations by piecewise smooth functions and associated variational problems. Communications on Pure and Applied Mathematics. 1989;42(5):577-685.

[154] Vardhanabhuti S, Jeng XJ, Wu Y, Li H. Parametric modeling of whole-genome sequencing data for CNV identification. Biostatistics. 2014;15(3):427-41.

[155] Wang LY, Abyzov A, Korbel JO, Snyder M, Gerstein M. MSB: a mean-shift-based approach for the analysis of structural variation in the genome. Genome Res. 2009;19(1):106-17.

[156] Yuan X, Yu J, Xi J, Yang L, Shang J, Li Z, et al. CNV_IFTV: An Isolation Forest and Total Variation-Based Detection of CNVs from Short-Read Sequencing Data. IEEE/ACM Trans Comput Biol Bioinform. 2021;18(2):539-49.

[157] Zhang Y, Liu W, Duan J. On the core segmentation algorithms of copy number variation detection tools. Briefings in Bioinformatics. 2024;25(2).

[158] Abyzov A, Li S, Kim DR, Mohiyuddin M, Stütz AM, Parrish NF, et al. Analysis of deletion breakpoints from 1,092 humans reveals details of mutation mechanisms. Nature Communications. 2015;6(1):7256.

[159] Babadi M, Fu JM, Lee SK, Smirnov AN, Gauthier LD, Walker M, et al. GATK-gCNV enables the discovery of rare copy number variants from exome sequencing data. Nat Genet. 2023;55(9):1589-97.

[160] Backenroth D, Homsy J, Murillo LR, Glessner J, Lin E, Brueckner M, et al. CANOES: detecting rare copy number variants from whole exome sequencing data. Nucleic Acids Res. 2014;42(12):e97.

[161] Bartenhagen C, Dugas M. Robust and exact structural variation detection with paired-end and soft-clipped alignments: SoftSV compared with eight algorithms. Briefings in Bioinformatics. 2015;17(1):51-62.

[162] Becker T, Lee W-P, Leone J, Zhu Q, Zhang C, Liu S, et al. FusorSV: an algorithm for optimally combining data from multiple structural variation detection methods. Genome Biology. 2018;19(1):38.

[163] Behera S, Catreux S, Rossi M, Truong S, Huang Z, Ruehle M, et al. Comprehensive genome analysis and variant detection at scale using DRAGEN. Nature Biotechnology. 2025;43(7):1177-91.

[164] Boeva V, Popova T, Bleakley K, Chiche P, Cappo J, Schleiermacher G, et al. Control-FREEC: a tool for assessing copy number and allelic content using next- generation sequencing data. Bioinformatics. 2011;28(3):423-5.

[165] Cameron DL, Schröder J, Penington JS, Do H, Molania R, Dobrovic A, et al. GRIDSS: sensitive and specific genomic rearrangement detection using positional de Bruijn graph assembly. Genome Res. 2017;27(12):2050-60.

[166] Chen K, Wallis JW, McLellan MD, Larson DE, Kalicki JM, Pohl CS, et al. BreakDancer: an algorithm for high-resolution mapping of genomic structural variation. Nat Methods. 2009;6(9):677-81.

[167] Chen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Källberg M, et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics. 2015;32(8):1220-2. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[168] Chiang T, Liu X, Wu TJ, Hu J, Sedlazeck FJ, White S, et al. Atlas-CNV: a validated approach to call single-exon CNVs in the eMERGESeq gene panel. Genet Med. 2019;21(9):2135-44.

[169] D’Aurizio R, Pippucci T, Tattini L, Giusti B, Pellegrini M, Magi A. Enhanced copy number variants detection from whole-exome sequencing data using EXCAVATOR2. Nucleic Acids Res. 2016;44(20):e154.

[170] Demidov G, Sturm M, Ossowski S. ClinCNV: multi-sample germline CNV detection in NGS data. bioRxiv. 2022:2022.06.10.495642.

[171] Dharanipragada P, Vogeti S, Parekh N. iCopyDAV: Integrated platform for copy number variations-Detection, annotation and visualization. PLoS One. 2018;13(4):e0195334.

[172] Dierckxsens N, Li T, Vermeesch JR, Xie Z. A benchmark of structural variation detection by long reads through a realistic simulated model. Genome Biol. 2021;22(1):342.

[173] Eisfeldt J, Vezzi F, Olason P, Nilsson D, Lindstrand A. TIDDIT, an efficient and comprehensive structural variant caller for massive parallel sequencing data. F1000Res. 2017;6:664.

[174] Fang L, Hu J, Wang D, Wang K. NextSV: a meta-caller for structural variants from low-coverage long-read sequencing data. BMC Bioinformatics. 2018;19(1):180.

[175] Fowler A. DECoN: A Detection and Visualization Tool for Exonic Copy Number Variants. Methods Mol Biol. 2022;2493:77-88.

[176] Gambin T, Akdemir ZC, Yuan B, Gu S, Chiang T, Carvalho CMB, et al. Homozygous and hemizygous CNV detection from exome sequencing data in a Mendelian disease cohort. Nucleic Acids Res. 2017;45(4):1633-48.

[177] Jeffares DC, Jolly C, Hoti M, Speed D, Shaw L, Rallis C, et al. Transient structural variations have strong effects on quantitative traits and reproductive isolation in fission yeast. Nat Commun. 2017;8:14061.

[178] Jiang Y, Oldridge DA, Diskin SJ, Zhang NR. CODEX: a normalization and copy number variation detection method for whole exome sequencing. Nucleic Acids Res. 2015;43(6):e39.

[179] Jiang Y, Wang R, Urrutia E, Anastopoulos IN, Nathanson KL, Zhang NR. CODEX2: full-spectrum copy number variation detection by high-throughput DNA sequencing. Genome Biology. 2018;19(1):202.

[180] Johansson LF, van Dijk F, de Boer EN, van Dijk-Bos KK, Jongbloed JD, van der Hout AH, et al. CoNVaDING: Single Exon Variation Detection in Targeted NGS Data. Hum Mutat. 2016;37(5):457-64.

[181] Klambauer G, Schwarzbauer K, Mayr A, Clevert DA, Mitterecker A, Bodenhofer U, et al. cn.MOPS: mixture of Poissons for discovering copy number variations in next- generation sequencing data with a low false discovery rate. Nucleic Acids Res. 2012;40(9):e69.

[182] Kronenberg ZN, Osborne EJ, Cone KR, Kennedy BJ, Domyan ET, Shapiro MD, et al. Wham: Identifying Structural Variants of Biological Consequence. PLoS Comput Biol. 2015;11(12):e1004572. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[183] Krumm N, Sudmant PH, Ko A, O’Roak BJ, Malig M, Coe BP, et al. Copy number variation detection and genotyping from exome sequence data. Genome Res. 2012;22(8):1525-32.

[184] Lam HY, Pan C, Clark MJ, Lacroute P, Chen R, Haraksingh R, et al. Detecting and annotating genetic variations using the HugeSeq pipeline. Nat Biotechnol. 2012;30(3):226-9.

[185] Lam HYK, Mu XJ, Stütz AM, Tanzer A, Cayting PD, Snyder M, et al. Nucleotide- resolution analysis of structural variants using BreakSeq and a breakpoint library. Nature Biotechnology. 2010;28(1):47-55.

[186] Layer RM, Chiang C, Quinlan AR, Hall IM. LUMPY: a probabilistic framework for structural variant discovery. Genome Biology. 2014;15(6):R84.

[187] Li J, Lupat R, Amarasinghe KC, Thompson ER, Doyle MA, Ryland GL, et al. CONTRA: copy number analysis for targeted resequencing. Bioinformatics. 2012;28(10):1307-13.

[188] Love MI, Myšičková A, Sun R, Kalscheuer V, Vingron M, Haas SA. Modeling read counts for CNV detection in exome sequencing data. Stat Appl Genet Mol Biol. 2011;10(1).

[189] Michaelson JJ, Sebat J. forestSV: structural variant discovery through statistical learning. Nature Methods. 2012;9(8):819-21.

[190] Mohiyuddin M, Mu JC, Li J, Bani Asadi N, Gerstein MB, Abyzov A, et al. MetaSV: an accurate and integrative structural-variant caller for next generation sequencing. Bioinformatics. 2015;31(16):2741-4.

[191] Packer JS, Maxwell EK, O’Dushlaine C, Lopez AE, Dewey FE, Chernomorsky R, et al. CLAMMS: a scalable algorithm for calling common and rare copy number variants from exome sequencing data. Bioinformatics. 2016;32(1):133-5.

[192] Plagnol V, Curtis J, Epstein M, Mok KY, Stebbings E, Grigoriadou S, et al. A robust model for read count data in exome sequencing experiments and implications for copy number variant calling. Bioinformatics. 2012;28(21):2747-54.

[193] Popic V, Rohlicek C, Cunial F, Hajirasouliha I, Meleshko D, Garimella K, et al. Cue: a deep-learning framework for structural variant discovery and genotyping. Nat Methods. 2023;20(4):559-68.

[194] Pounraja VK, Jayakar G, Jensen M, Kelkar N, Girirajan S. A machine-learning approach for accurate detection of copy number variants from exome sequencing. Genome Res. 2019;29(7):1134-43.

[195] Povysil G, Tzika A, Vogt J, Haunschmid V, Messiaen L, Zschocke J, et al. panelcn.MOPS: Copy-number detection in targeted NGS panel data for clinical diagnostics. Hum Mutat. 2017;38(7):889-97.

[196] Rausch T, Zichner T, Schlattl A, Stütz AM, Benes V, Korbel JO. DELLY: structural variant discovery by integrated paired-end and split-read analysis. Bioinformatics. 2012;28(18):i333-i9.

[197] Roller E, Ivakhno S, Lee S, Royce T, Tanner S. Canvas: versatile and scalable detection of copy number variants. Bioinformatics. 2016;32(15):2375-7.

[198] Samarakoon PS, Sorte HS, Kristiansen BE, Skodje T, Sheng Y, Tjønnfjord GE, et al. Identification of copy number variants from exome sequence data. BMC Genomics. 2014;15(1):661. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[199] Sathirapongsasuti JF, Lee H, Horst BA, Brunner G, Cochran AJ, Binder S, et al. Exome sequencing-based copy-number variation and loss of heterozygosity detection: ExomeCNV. Bioinformatics. 2011;27(19):2648-54.

[200] Talevich E, Shain AH, Botton T, Bastian BC. CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLOS Computational Biology. 2016;12(4):e1004873.

[201] Wang C, Evans JM, Bhagwate AV, Prodduturi N, Sarangi V, Middha M, et al. PatternCNV: a versatile tool for detecting copy number changes from exome sequencing data. Bioinformatics. 2014;30(18):2678-80.

[202] Wong K, Keane TM, Stalker J, Adams DJ. Enhanced structural variant and breakpoint detection using SVMerge by integration of multiple detection methods and local assembly. Genome Biology. 2010;11(12):R128.

[203] Xi R, Lee S, Xia Y, Kim TM, Park PJ. Copy number analysis of whole-genome data using BIC-seq2 and its application to detection of cancer susceptibility variants. Nucleic Acids Res. 2016;44(13):6274-86.

[204] Ye K, Schulz MH, Long Q, Apweiler R, Ning Z. Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads. Bioinformatics. 2009;25(21):2865-71.

[205] Yoon S, Xuan Z, Makarov V, Ye K, Sebat J. Sensitive and accurate detection of copy number variants using read depth of coverage. Genome Res. 2009;19(9):1586-92.

[206] Zarate S, Carroll A, Mahmoud M, Krasheninina O, Jun G, Salerno WJ, et al. Parliament2: Accurate structural variant calling at scale. Gigascience. 2020;9(12).

[207] Zhang J, Wang J, Wu Y. An improved approach for accurate and efficient calling of structural variations with low-coverage sequence data. BMC Bioinformatics. 2012;13 Suppl 6(Suppl 6):S6.

[208] Zhu M, Need AC, Han Y, Ge D, Maia JM, Zhu Q, et al. Using ERDS to infer copy- number variants in high-coverage genomes. Am J Hum Genet. 2012;91(3):408-21.

[209] Litmaps. Litmaps (Version 2025-01-16) [Visualization purposes]. 2024.

[210] Becker T, Lee WP, Leone J, Zhu Q, Zhang C, Liu S, et al. FusorSV: an algorithm for optimally combining data from multiple structural variation detection methods. Genome Biol. 2018;19(1):38.

[211] Majidian S, Agustinho DP, Chin C-S, Sedlazeck FJ, Mahmoud M. Genomic variant benchmark: if you cannot measure it, you cannot improve it. Genome Biology. 2023;24(1):221.

[212] Zook JM, Hansen NF, Olson ND, Chapman L, Mullikin JC, Xiao C, et al. A robust benchmark for detection of germline large deletions and insertions. Nature Biotechnology. 2020;38(11):1347-55.

[213] Kronenberg Z, Nolan C, Porubsky D, Mokveld T, Rowell WJ, Lee S, et al. The Platinum Pedigree: a long-read benchmark for genetic variants. Nature Methods. 2025;22(8):1669-76.

[214] Narcı K, Vannieuwkerke N, Garcia MU, Bot N-C, EladH, Kirk J, et al. Nf- core/variantbenchmarking: 1.2.0 doubtful Adams. Zenodo; 2025.

[215] Hukku A, Pividori M, Luca F, Pique-Regi R, Im HK, Wen X. Probabilistic colocalization of genetic variants from complex and molecular traits: promise and limitations. The american journal of human genetics. 2021;108(1):25-35. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[216] English AC, Menon VK, Gibbs RA, Metcalf GA, Sedlazeck FJ. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biology. 2022;23(1):271.

[217] Yang J, Chaisson MJP. TT-Mars: structural variants assessment based on haplotype-resolved assemblies. Genome Biology. 2022;23(1):110.

[218] Wagner J, Olson ND, Harris L, McDaniel J, Cheng H, Fungtammasan A, et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat Biotechnol. 2022;40(5):672-80.

[219] Olson ND, Wagner J, McDaniel J, Stephens SH, Westreich ST, Prasanna AG, et al. PrecisionFDA Truth Challenge V2: Calling variants from short and long reads in difficult- to-map regions. Cell Genomics. 2022;2(5).

[220] Moreno-Cabrera JM, Del Valle J, Castellanos E, Feliubadaló L, Pineda M, Brunet J, et al. Evaluation of CNV detection tools for NGS panel data in genetic diagnostics. Eur J Hum Genet. 2020;28(12):1645-55.

[221] Parikh H, Mohiyuddin M, Lam HY, Iyer H, Chen D, Pratt M, et al. svclassify: a method to establish benchmark structural variant calls. BMC Genomics. 2016;17:64.

[222] Carter NP. Methods and strategies for analyzing copy number variation using DNA microarrays. Nat Genet. 2007;39(7 Suppl):S16-21.

[223] Lepkes L, Kayali M, Blümcke B, Weber J, Suszynska M, Schmidt S, et al. Performance of In Silico Prediction Tools for the Detection of Germline Copy Number Variations in Cancer Predisposition Genes in 4208 Female Index Patients with Familial Breast and Ovarian Cancer. Cancers (Basel). 2021;13(1).

[224] Nutsua ME, Fischer A, Nebel A, Hofmann S, Schreiber S, Krawczak M, et al. Family-Based Benchmarking of Copy Number Variation Detection Software. PLoS One. 2015;10(7):e0133465.

[225] Roca I, González-Castro L, Fernández H, Couce ML, Fernández-Marmiesse A. Free-access copy-number variant detection tools for targeted next-generation sequencing data. Mutat Res Rev Mutat Res. 2019;779:114-25.

[226] Royer-Bertrand B, Cisarova K, Niel-Butschi F, Mittaz-Crettol L, Fodstad H, Superti-Furga A. CNV Detection from Exome Sequencing Data in Routine Diagnostics of Rare Genetic Disorders: Opportunities and Limitations. Genes (Basel). 2021;12(9).

[227] Silva C, Ferrão J, Marques B, Pedro S, Correia H, Valente A, et al. Comparative analysis of hybrid-SNP microarray and nanopore sequencing for detection of large- sized copy number variants in the human genome. Mol Cytogenet. 2025;18(1):18.

[228] Smolander J, Khan S, Singaravelu K, Kauko L, Lund RJ, Laiho A, et al. Evaluation of tools for identifying large copy number variations from ultra-low-coverage whole- genome sequencing data. BMC Genomics. 2021;22(1):357.

[229] Tan R, Wang Y, Kleinstein SE, Liu Y, Zhu X, Guo H, et al. An evaluation of copy number variation detection tools from whole-exome sequencing data. Hum Mutat. 2014;35(7):899-907.

[230] Zhang L, Bai W, Yuan N, Du Z. Comprehensively benchmarking applications for detecting copy number variation. PLoS Comput Biol. 2019;15(5):e1007069.

[231] Zhao S, Xu D, Cai J, Shen Q, He M, Pan X, et al. Benchmarking strategies for CNV calling from whole genome bisulfite data in humans. Comput Struct Biotechnol J. 2025;27:912-9. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[232] Zhou A, Lin T, Xing J. Evaluating nanopore sequencing data processing pipelines for structural variation identification. Genome Biol. 2019;20(1):237.

[233] Handsaker RE, Korn JM, Nemesh J, McCarroll SA. Discovery and genotyping of genome structural polymorphism by sequencing on a population scale. Nat Genet. 2011;43(3):269-76.

[234] Montanucci L, Lewis-Smith D, Collins RL, Niestroj LM, Parthasarathy S, Xian J, et al. Genome-wide identification and phenotypic characterization of seizure-associated copy number variations in 741,075 individuals. Nat Commun. 2023;14(1):4392.

[235] Saarentaus EC, Havulinna AS, Mars N, Ahola-Olli A, Kiiskinen TTJ, Partanen J, et al. Polygenic burden has broader impact on health, cognition, and socioeconomic outcomes than most rare and high-risk copy number variants. Mol Psychiatry. 2021;26(9):4884-95.

[236] Warland A, Kendall KM, Rees E, Kirov G, Caseras X. Schizophrenia-associated genomic copy number variants and subcortical brain volumes in the UK Biobank. Mol Psychiatry. 2020;25(4):854-62.

[237] Kendall KM, Bracher-Smith M, Fitzpatrick H, Lynham A, Rees E, Escott-Price V, et al. Cognitive performance and functional outcomes of carriers of pathogenic copy number variants: analysis of the UK Biobank. Br J Psychiatry. 2019;214(5):297-304.

[238] Lal D, Ruppert AK, Trucks H, Schulz H, de Kovel CG, Kasteleijn-Nolst Trenité D, et al. Burden analysis of rare microdeletions suggests a strong impact of neurodevelopmental genes in genetic generalised epilepsies. PLoS Genet. 2015;11(5):e1005226.

[239] Simpson NH, Ceroni F, Reader RH, Covill LE, Knight JC, Hennessy ER, et al. Genome-wide analysis identifies a role for common copy number variants in specific language impairment. Eur J Hum Genet. 2015;23(10):1370-7.

[240] Sokolowski M, Wasserman J, Wasserman D. Rare CNVs in Suicide Attempt include Schizophrenia-Associated Loci and Neurodevelopmental Genes: A Pilot Genome-Wide and Family-Based Study. PLoS One. 2016;11(12):e0168531.

[241] Vevera J, Zarrei M, Hartmannová H, Jedličková I, Mušálková D, Přistoupilová A, et al. Rare copy number variation in extremely impulsively violent males. Genes Brain Behav. 2019;18(6):e12536.

[242] Green EK, Rees E, Walters JT, Smith KG, Forty L, Grozeva D, et al. Copy number variation in bipolar disorder. Mol Psychiatry. 2016;21(1):89-93.

[243] Sinnott-Armstrong N, Tanigawa Y, Amar D, Mars N, Benner C, Aguirre M, et al. Genetics of 35 blood and urine biomarkers in the UK Biobank. Nat Genet. 2021;53(2):185-94.

[244] Leppa VM, Kravitz SN, Martin CL, Andrieux J, Le Caignec C, Martin-Coignard D, et al. Rare Inherited and De Novo CNVs Reveal Complex Contributions to ASD Risk in Multiplex Families. Am J Hum Genet. 2016;99(3):540-54.

[245] Monlong J, Girard SL, Meloche C, Cadieux-Dion M, Andrade DM, Lafreniere RG, et al. Global characterization of copy number variants in epilepsy patients from whole genome sequencing. PLoS Genet. 2018;14(4):e1007285.

[246] Chen J, Calhoun VD, Perrone-Bizzozero NI, Pearlson GD, Sui J, Du Y, et al. A pilot study on commonality and specificity of copy number variants in schizophrenia and bipolar disorder. Transl Psychiatry. 2016;6(5):e824. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[247] Chen L, Abel HJ, Das I, Larson DE, Ganel L, Kanchi KL, et al. Association of structural variation with cardiometabolic traits in Finns. Am J Hum Genet. 2021;108(4):583-96.

[248] de Jesús Ascencio-Montiel I, Pinto D, Parra EJ, Valladares-Salgado A, Cruz M, Scherer SW. Characterization of Large Copy Number Variation in Mexican Type 2 Diabetes subjects. Sci Rep. 2017;7(1):17105.

[249] Yin CL, Chen HI, Li LH, Chien YL, Liao HM, Chou MC, et al. Genome-wide analysis of copy number variations identifies PARK2 as a candidate gene for autism spectrum disorder. Mol Autism. 2016;7:23.

[250] Wu MC, Lee S, Cai T, Li Y, Boehnke M, Lin X. Rare-variant association testing for sequencing data with the sequence kernel association test. Am J Hum Genet. 2011;89(1):82-93.

[251] Lee S, Emond MJ, Bamshad MJ, Barnes KC, Rieder MJ, Nickerson DA, et al. Optimal unified approach for rare-variant association testing with application to small- sample case-control whole-exome sequencing studies. Am J Hum Genet. 2012;91(2):224-37.

[252] Zhan X, Girirajan S, Zhao N, Wu MC, Ghosh D. A novel copy number variants kernel association test with application to autism spectrum disorders studies. Bioinformatics. 2016;32(23):3603-10.

[253] Brucker A, Lu W, Marceau West R, Yu QY, Hsiao CK, Hsiao TH, et al. Association test using Copy Number Profile Curves (CONCUR) enhances power in rare copy number variant analysis. PLoS Comput Biol. 2020;16(5):e1007797.

[254] Maus Esfahani N, Catchpoole D, Khan J, Kennedy PJ. MCKAT: a multi- dimensional copy number variant kernel association test. BMC Bioinformatics. 2021;22(1):588.

[255] Maus Esfahani N, Catchpoole D, Kennedy PJ. SMCKAT, a Sequential Multi- Dimensional CNV Kernel-Based Association Test. Life (Basel). 2021;11(12).

[256] Kang HM, Sul JH, Service SK, Zaitlen NA, Kong S-y, Freimer NB, et al. Variance component model to account for sample structure in genome-wide association studies. Nature Genetics. 2010;42(4):348-54.

[257] Loh PR, Tucker G, Bulik-Sullivan BK, Vilhjálmsson BJ, Finucane HK, Salem RM, et al. Efficient Bayesian mixed-model analysis increases association power in large cohorts. Nat Genet. 2015;47(3):284-90.

[258] Loh PR, Kichaev G, Gazal S, Schoech AP, Price AL. Mixed-model association for biobank-scale datasets. Nat Genet. 2018;50(7):906-8.

[259] Ionita-Laza I, Perry GH, Raby BA, Klanderman B, Lee C, Laird NM, et al. On the analysis of copy-number variations in genome-wide association studies: a translation of the family-based association test. Genetic Epidemiology. 2008;32(3):273-84.

[260] Liu M, Moon S, Wang L, Kim S, Kim YJ, Hwang MY, et al. On the association analysis of CNV data: a fast and robust family-based association method. BMC Bioinformatics. 2017;18(1):217.

[261] Wu X, Huai C, Shen L, Li M, Yang C, Zhang J, et al. Genome-wide study of copy number variation implicates multiple novel loci for schizophrenia risk in Han Chinese family trios. iScience. 2021;24(8). ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[262] Mendes M, Chen DZ, Engchuan W, Leal TP, Thiruvahindrapuram B, Trost B, et al. Chromosome X-wide common variant association study in autism spectrum disorder. Am J Hum Genet. 2025;112(1):135-53.

[263] Han J, Walters JT, Kirov G, Pocklington A, Escott-Price V, Owen MJ, et al. Gender differences in CNV burden do not confound schizophrenia CNV associations. Sci Rep. 2016;6:25986.

[264] Subirana I, Diaz-Uriarte R, Lucas G, Gonzalez JR. CNVassoc: Association analysis of CNV data using R. BMC Med Genomics. 2011;4:47.

[265] Wittig M, Helbig I, Schreiber S, Franke A. CNVineta: a data mining tool for large case-control copy number variation datasets. Bioinformatics. 2010;26(17):2208-9.

[266] Prakash SK, Bondy CA, Maslen CL, Silberbach M, Lin AE, Perrone L, et al. Autosomal and X chromosome structural variants are associated with congenital heart defects in Turner syndrome: The NHLBI GenTAC registry. Am J Med Genet A. 2016;170(12):3157-64.

[267] Wang H, Dombroski BA, Cheng PL, Tucci A, Si YQ, Farrell JJ, et al. Structural variation detection and association analysis of whole-genome-sequence data from 16,543 Alzheimer’s disease sequencing project subjects. Alzheimers Dement. 2025;21(6):e70277.

[268] Jeng J, Wu Q, Li H. A Statistical Method for Identifying Trait-Associated Copy Number Variants. Hum Hered. 2015;79(3-4):147-56.

[269] Labani M, Afrasiabi A, Beheshti A, Lovell NH, Alinejad-Rokny H. PeakCNV: A multi-feature ranking algorithm-based tool for genome-wide copy number variation- association study. Comput Struct Biotechnol J. 2022;20:4975-83.

[270] Alinejad-Rokny H, Heng JIT, Forrest ARR. Brain-Enriched Coding and Long Non- coding RNA Genes Are Overrepresented in Recurrent Neurodevelopmental Disorder CNVs. Cell Rep. 2020;33(4):108307.

[271] Larsen SJ, do Canto LM, Rogatto SR, Baumbach J. CoNVaQ: a web tool for copy number variation-based association studies. BMC Genomics. 2018;19(1):369.

[272] Abe-Hatano C, Iida A, Kosugi S, Momozawa Y, Terao C, Ishikawa K, et al. Whole genome sequencing of 45 Japanese patients with intellectual disability. Am J Med Genet A. 2021;185(5):1468-80.

[273] Balagué-Dobón L, Cáceres A, González JR. Fully exploiting SNP arrays: a systematic review on the tools to extract underlying genomic structure. Brief Bioinform. 2022;23(2).

[274] Billingsley KJ, Ding J, Jerez PA, Illarionova A, Levine K, Grenn FP, et al. Genome- Wide Analysis of Structural Variants in Parkinson Disease. Ann Neurol. 2023;93(5):1012-22.

[275] de Almeida Santana MH, Junior GA, Cesar AS, Freua MC, da Costa Gomes R, da Luz ESS, et al. Copy number variations and genome-wide associations reveal putative genes and metabolic pathways involved with the feed conversion ratio in beef cattle. J Appl Genet. 2016;57(4):495-504.

[276] de Lemos MVA, Peripolli E, Berton MP, Feitosa FLB, Olivieri BF, Stafuzza NB, et al. Association study between copy number variation and beef fatty acid profile of Nellore cattle. J Appl Genet. 2018;59(2):203-23. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[277] Fernandes AC, da Silva VH, Goes CP, Moreira GCM, Godoy TF, Ibelli AMG, et al. Genome-wide detection of CNVs and their association with performance traits in broilers. BMC Genomics. 2021;22(1):354.

[278] Getmantseva L, Kolosova M, Fede K, Korobeinikova A, Kolosov A, Romanets E, et al. Finding Predictors of Leg Defects in Pigs Using CNV-GWAS. Genes (Basel). 2023;14(11).

[279] Glessner JT, Khan ME, Chang X, Liu Y, Otieno FG, Lemma M, et al. Rare recurrent copy number variations in metabotropic glutamate receptor interacting genes in children with neurodevelopmental disorders. J Neurodev Disord. 2023;15(1):14.

[280] Glessner JT, Li J, Wang D, March M, Lima L, Desai A, et al. Copy number variation meta-analysis reveals a novel duplication at 9p24 associated with multiple neurodevelopmental disorders. Genome Med. 2017;9(1):106.

[281] Han N, Oh JM, Kim IW. Combination of Genome-Wide Polymorphisms and Copy Number Variations of Pharmacogenes in Koreans. J Pers Med. 2021;11(1).

[282] Irvin MR, Wineinger NE, Rice TK, Pajewski NM, Kabagambe EK, Gu CC, et al. Genome-Wide Detection of Allele Specific Copy Number Variation Associated with Insulin Resistance in African Americans from the HyperGEN Study. PLOS ONE. 2011;6(8):e24052.

[283] Kaivola K, Chia R, Ding J, Rasheed M, Fujita M, Menon V, et al. Genome-wide structural variant analysis identifies risk loci for non-Alzheimer’s dementias. Cell Genom. 2023;3(6):100316.

[284] Kasak L, Rull K, Sõber S, Laan M. Copy number variation profile in the placental and parental genomes of recurrent pregnancy loss families. Sci Rep. 2017;7:45327.

[285] León LE, Benavides F, Espinoza K, Vial C, Alvarez P, Palomares M, et al. Partial microduplication in the histone acetyltransferase complex member KANSL1 is associated with congenital heart defects in 22q11.2 microdeletion syndrome patients. Sci Rep. 2017;7(1):1795.

[286] Li D, Matsuoka LS, Donoghue S, Hou C, Strong A, McDonald-McGinn DM, et al. Modeling the long-range effect of an inversion downstream of EFNB1 concludes a 43- year molecular diagnostic odyssey for craniofrontonasal syndrome. Eur J Hum Genet. 2025.

[287] Oliveira P, Costa GNO, Damasceno AKA, Hartwig FP, Barbosa GCG, Figueiredo CA, et al. Genome-wide burden and association analyses implicate copy number variations in asthma risk among children and young adults from Latin America. Sci Rep. 2018;8(1):14475.

[288] Pathak GA, Polimanti R, Silzer TK, Wendt FR, Chakraborty R, Phillips NR. Genetically-regulated transcriptomics & copy number variation of proctitis points to altered mitochondrial and DNA repair mechanisms in individuals of European ancestry. BMC Cancer. 2020;20(1):954.

[289] Rambo-Martin BL, Mulle JG, Cutler DJ, Bean LJH, Rosser TC, Dooley KJ, et al. Analysis of Copy Number Variants on Chromosome 21 in Down Syndrome-Associated Congenital Heart Defects. G3 (Bethesda). 2018;8(1):105-11.

[290] Rymuza J, Kober P, Maksymowicz M, Nyc A, Mossakowska BJ, Woroniecka R, et al. High level of aneuploidy and recurrent loss of chromosome 11 as relevant features of somatotroph pituitary tumors. J Transl Med. 2024;22(1):994. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[291] Sasaki S, Watanabe T, Ibi T, Hasegawa K, Sakamoto Y, Moriwaki S, et al. Identification of deleterious recessive haplotypes and candidate deleterious recessive mutations in Japanese Black cattle. Sci Rep. 2021;11(1):6687.

[292] Sha Z, Sun KY, Jung B, Barzilay R, Moore TM, Almasy L, et al. Copy Number Variant Architecture of Child Psychopathology and Cognitive Development in the ABCD Study. Am J Psychiatry. 2025;182(8):763-78.

[293] Silva VH, Regitano LC, Geistlinger L, Pértille F, Giachetto PF, Brassaloti RA, et al. Genome-Wide Detection of CNVs and Their Association with Meat Tenderness in Nelore Cattle. PLoS One. 2016;11(6):e0157711.

[294] Wang K, Cadzow M, Bixley M, Leask MP, Merriman ME, Yang Q, et al. A Polynesian-specific copy number variant encompassing the MICA gene associates with gout. Hum Mol Genet. 2022;31(21):3757-68.

[295] Wu J, Wu T, Xie X, Niu Q, Zhao Z, Zhu B, et al. Genetic Association Analysis of Copy Number Variations for Meat Quality in Beef Cattle. Foods. 2023;12(21).

[296] Wu Y, Adams K. Genome-Wide Analyses of Copy Number Variants in 751 Populus trichocarpa Individuals From Natural Populations. Genome Biol Evol. 2025;17(7).

[297] Xie X, Shi L, Hou G, Zhong Z, Wang Z, Pan D, et al. Genome wide detection of CNV and their association with body size in Danzhou chickens. Poult Sci. 2024;103(12):104266.

[298] Zhang M, Li Q, Wang KL, Dong Y, Mu YT, Cao YM, et al. Lipolysis and gestational diabetes mellitus onset: a case-cohort genome-wide association study in Chinese. J Transl Med. 2023;21(1):47.

[299] Zhang W, Yang P, Yang Y, Liu S, Xu Y, Wu C, et al. Genomic landscape and distinct molecular subtypes of primary testicular lymphoma. J Transl Med. 2024;22(1):414.

[300] Zhang X, Du R, Li S, Zhang F, Jin L, Wang H. Evaluation of copy number variation detection for a SNP array platform. BMC Bioinformatics. 2014;15(1):50.

[301] Nurk S, Koren S, Rhie A, Rautiainen M, Bzikadze AV, Mikheenko A, et al. The complete sequence of a human genome. Science. 2022;376(6588):44-53.

[302] Eizenga JM, Novak AM, Sibbesen JA, Heumos S, Ghaffaari A, Hickey G, et al. Pangenome Graphs. Annu Rev Genomics Hum Genet. 2020;21:139-62.

[303] Chin CS, Behera S, Khalak A, Sedlazeck FJ, Sudmant PH, Wagner J, et al. Multiscale analysis of pangenomes enables improved representation of genomic diversity for repetitive and clinically relevant genes. Nat Methods. 2023;20(8):1213-21.

[304] Groza C, Schwendinger-Schreck C, Cheung WA, Farrow EG, Thiffault I, Lake J, et al. Pangenome graphs improve the analysis of structural variants in rare genetic diseases. Nat Commun. 2024;15(1):657.

[305] Ebler J, Ebert P, Clarke WE, Rausch T, Audano PA, Houwaart T, et al. Pangenome- based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes. Nat Genet. 2022;54(4):518-25.

[306] Liao W-W, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, et al. A draft human pangenome reference. Nature. 2023;617(7960):312-24.

[307] Eggertsson HP, Kristmundsdottir S, Beyter D, Jonsson H, Skuladottir A, Hardarson MT, et al. GraphTyper2 enables population-scale genotyping of structural variation using pangenome graphs. Nat Commun. 2019;10(1):5402. ACCEPTED MANUSCRIPT ARTICLE IN PRESS ARTICLE IN PRESS

[308] Eggertsson HP, Jonsson H, Kristmundsdottir S, Hjartarson E, Kehr B, Masson G, et al. Graphtyper enables population-scale genotyping using pangenome graphs. Nat Genet. 2017;49(11):1654-60.

[309] Bai W-Y, Liu S, Duan Z, Yang J-J, Chen J, Hou J, et al. Genome-wide associations of structural variants with human traits through imputation from long-read assemblies. Nature Genetics. 2026;58(6):1258-67.

[310] Bergen SE, Ploner A, Howrigan D, O’Donovan MC, Smoller JW, Sullivan PF, et al. Joint Contributions of Rare Copy Number Variants and Common SNPs to Risk for Schizophrenia. Am J Psychiatry. 2019;176(1):29-35.

[311] Gillani R, Collins RL, Crowdis J, Garza A, Jones JK, Walker M, et al. Rare germline structural variants increase risk for pediatric solid tumors. Science. 2025;387(6729):eadq0071.

[312] Ewels PA, Peltzer A, Fillinger S, Patel H, Alneberg J, Wilm A, et al. The nf-core framework for community-curated bioinformatics pipelines. Nat Biotechnol. 2020;38(3):276-8.

[313] Porcu E, Rüeger S, Lepik K, Santoni FA, Reymond A, Kutalik Z. Mendelian randomization integrating GWAS and eQTL data reveals genetic determinants of complex and clinical traits. Nat Commun. 2019;10(1):3300.

[314] Birney E, Vamathevan J, Goodhand P. Genomics in healthcare: GA4GH looks to 2022. bioRxiv. 2017:203554.

[315] Hofmeister RJ, Ribeiro DM, Rubinacci S, Delaneau O. Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank. Nat Genet. 2023;55(7):1243-9.

[316] Collins RL, Glessner JT, Porcu E, Lepamets M, Brandon R, Lauricella C, Han L, Morley T, Niestroj LM, Ulirsch J, Everett S, Howrigan DP, Boone PM, Fu J, Karczewski KJ, Kellaris G, Lowther C, Lucente D, Mohajeri K, Nõukas M, Nuttle X, Samocha KE, Gusella JF, Finucane H, Matyakhina L, Aradhya S, Meck J, Lal D, Neale BM, Hodge JC, Reymond A, Kutalik Z, Katsanis N, Davis EE, Hakonarson H, Sunyaev S, Brand H, Talkowski ME. A cross-disorder dosage sensitivity map of the human genome. Datasets. Zenodo. https://doi.org/10.5281/zenodo.6347673 (2022).

[317] Harris L, McDonagh EM, Zhang X, Fawcett K, Foreman A, Daneck P, Sergouniotis PI, Parkinson H, Mazzarotto F, Inouye M, Hollox EJ, Birney E, Fitzgerald T. Genome-wide association testing beyond SNPs. Supplementary Information 2. Datasets. Springer Nature. https://static-content.springer.com/esm/art%3A10.1038%2Fs41576-024- 00778-y/MediaObjects/41576_2024_778_MOESM2_ESM.xlsx (2025).

附加文件3:

Note S1. CNV检测的平台与方法

Note S2. 跨阵列与测序平台的CNV检测策略概述

短读长测序数据

        从短读长测序数据检测CNV通常采用五种主要策略。

        第一种是双端定位(paired-end mapping)或读长对方法。该策略检查双端或mate-pair读长之间的方向和间距。基于读长对的工具使用两种策略:基于模型和聚类。基于模型的策略使用概率检验识别读长对之间的异常距离,而聚类方法使用预定距离检测不一致的读数。

        第二种方法,分裂读长(split read),检查一个读长唯一比对而另一个无法比对或仅部分比对的读长对。分裂读长检测策略旨在识别包含结构变异(SVs)断点的读长,这些读长可能被分割成多个片段。部分比对的读长被分裂读长工具分解为若干片段,其中起始和末尾片段被独立重新比对。

        第三种策略涉及读长深度(read depth),通过测量比对到特定基因组区间的测序读长数量来检测CNV。基于读长深度的方法通常在固定大小或可变大小的不重叠基因组窗口中计算读长计数。比对、归一化、分段和拷贝数估计是关键阶段。在比对阶段,短读长被比对到参考基因组,通过计数以足够质量比对到每个基因组窗口的读长数量获得读长深度谱。在归一化阶段,校正由GC含量、序列可映射性或重复元件变异引起的读长深度系统偏倚以提高准确性。在分段阶段,每个个体的基因组被划分为表现出相似归一化读长深度的连续候选CNV区段。通常应用不同算法检测与拷贝数改变对应的变化点。最后,进行拷贝数状态估计,根据归一化读长深度谱为每个区段分配拷贝数值,从而识别受缺失或重复影响的基因组区域。从数学角度看,值得指出的是,一旦完成比对和归一化,从下一代测序(NGS)研究获得的读长深度数据与aCGH数据得出的探针log比值表现出相似性。因此,识别aCGH数据中CNV区域的传统方法可被改编用于NGS读长深度数据。这一步骤至关重要,因为它通过识别这些可被分类为扩增或缺失的区段,显著提高了CNV检测的准确性。已开发出若干用于CNV检测的分段算法和方法(见注释S3)。

        第四种方法是从头组装(de novo assembly, AS),它不依赖参考基因组,从重叠读长重建重叠群(contigs)。通过将重叠群比对到参考序列可推断CNV区段。AS方法为在广泛大小范围内检测新变异提供了无偏平台,尽管计算强度大且在重复区域表现不佳(8)。

        最后,第五种方法是组合方法(combinatorial approach),它整合多种类型的证据以提高CNV检出的灵敏度、特异性和分辨率。

长读长测序数据

        使用长读长测序数据进行包括CNV在内的SV检测依赖于多种计算策略,这些策略利用长读长的独特优势,如其跨越重复区域的能力和提供长距离信息的能力。主要方法之一是基于比对(alignment-based)的检测,即单个长读长被比对到参考基因组,并提取读长深度变化、分裂读长、软裁剪碱基和大比对缺口等特征。然后跨读长对这些特征进行聚类以形成共识CNV检出。cuteSV、Sniffles、Sniffles2、PBHoney、NanoSV、NanoVar和Picky等工具体现了这一策略,该方法通常快速,即使在相对较低的测序覆盖度下也表现良好。然而,结果可能对比对器的选择和测序错误谱敏感,并且可能难以解析重复区域中的复杂CNV。

        另一种广泛使用的策略是基于组装(assembly-based)的CNV检测。在该方法中,长读长首先通过从头或参考引导组装重建为重叠群。然后将这些重叠群与参考基因组比对以识别差异。SVIM-asm、SyRI和pbsv等工具遵循这种方法。基于组装的技术在断点检测方面提供高准确性,尤其适用于识别大片段插入、缺失和复杂重排,特别是在重复区域。然而,它们计算强度大,且通常需要更高的测序深度(通常高于20×)才能有效工作。

Note S3. 用于CNV检测的分段算法与方法

Note S4. CNV检测工具的平台特异性基准测试

Note S5. 性染色体CNV分析的挑战与考量

附加文件1:

表S2. 常见CNV检出器(基于高通量测序)

表S2. 常见CNV检出器(靶向panel、长读长测序与集成/合并方法)

        伯科靶向捕获Gene Panel基于直接合成的120nt 单链DNA探针和高性能捕获体系,覆盖均一性优异,批次稳定性极佳,不但可以为CNV分析提供稳定、均一的覆盖Reads,还可以根据个性化需求极速定制骨架探针模块,提高CNV检测性能,保证CNV的准确分析。

  

TargetCap® Core Exome Panel v3.0

       TargetCap@ Core Exome Panel v3.0基于伯科高品质DNA探针合成技术开发,全流程国产制造,由~40万条探针组成,以GRCh38/hg38人类参考基因组设计,参考Refseq、CCDS、ClinVar等数据库,覆盖19,524个基因,目标区域为33.9Mb。

捕获性能比较

伯科全外芯片v3.0性能优异与国外友商同类型产品v2相当,中靶率、覆盖率、覆盖均一性等参数均达到国际领先水平。

适配高通量流程平台

批次稳定

使用不同批次TargetCap® Core Exome Panel v3.0芯片对NA12878 gDNA进行捕获测序,结果显示,不同批次芯片在不同测序平台上均显示出优异的稳定性,不同位点的相对深度相关性高,批次稳定。

变异检测准确

单核苷酸变异(SNV)和插入缺失 (INDEL)是基因组变异的常见形式,也是引起人类疾病的重要原因。

 

选取NA12878标准品,与预期SNV和INDEL变异进行比较。结果表明,在MGI与Illumina测序平台,SNP灵敏度为99.1%,INDEL灵敏度为91.6%。

添加线粒体模块临床样本表现

20例全血样本(S1#-S20#),采用1-4 Plex方式使用TargetCap® Core Exome Panel v3.0添加线粒体模块进行过夜杂交捕获;其中,S1#-S10#在Illumina NovaSeq X Plus平台测序, S11#-S20#在MGI MGISEQ-T7平台测序,均采用150PE模式测序。得到测序数据后,抽取8Gb数据进行生信分析。

 

两种测序平台的数据表现相近,平均深度分别为111x/115x (Illumina/MGI),中靶率优异均> 85%,覆盖均一性极佳(0.2X_MD≥99.3%);仅使用8Gb数据,高达98.5%的捕获区域达到了30X以上,99.5%的捕获区域达到20X以上,为临床样本检测提供了可靠的捕获数据。

≥0.2X/0.5X_MD: Mean Depth,覆盖深度≥平均深度的0.2/0.5倍深度的区域占总区域的比例,用于表征覆盖均一性性能,越接近100%越好

 

 

TargetCap® Core Exome Panel v7.0

        TargetCap® Core Exome Panel v7.0(下文简称BOKE v7.0),该WES Panel增强了基因组hg19传统研究区域的覆盖,兼顾 hg19 & hg38 双版本基因组,可以更好的保证临床科研与转化的延续性。目标区域和捕获区域大小分别为40Mb和49Mb,对 hg19 传统研究区域覆盖提升至99.7%(友商I-v1),hg38传统研究区域覆盖相近(友商A-v8)。同时,新添加数百个具有一定功能与表型的基因,总基因数量达到20000+。

BOKE Core Exome Panel v7.0目标区域大小以及对不同友商产品目标区域的覆盖情况

 

在捕获性能方面,TargetCap® Core Exome Panel v7.0依然表现优异,与TargetCap® Core Exome Panel v3.0表现相近。在测序9Gb条件下,平均深度达到110x左右,20x和30x以上区域占比分别为99.5%和98.5%,Fold 80为1.5-1.6之间,与国际领先产品数据表现相当。

        同时,TargetCap® Core Exome Panel v7.0也可灵活的与拓展模块组合使用,满足不同场景的临床研究的需求及转化应用,包括线粒体、遗传病非编码区变异位点、单基因全覆盖、病毒基因组、肿瘤全景变异检测、重大疾病多基因风险评估模块等。此外,伯科公司自研自造的寡核苷酸合成平台可以快速响应个性化定制的需求,为人类基因组分子遗传学的研究与转化,提供更加全面高效的解决方案。

推荐阅读