Repository navigation
Expand file tree
/
Copy pathmain.tex
More file actions
574 lines (418 loc) · 60.7 KB
/
Copy pathmain.tex
File metadata and controls
574 lines (418 loc) · 60.7 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% Please note that whilst this template provides a
% preview of the typeset manuscript for submission, it
% will not necessarily be the final publication layout.
%
% letterpaper/a4paper: US/UK paper size toggle
% num-refs/alpha-refs: numeric/author-year citation and bibliography toggle
%\documentclass[letterpaper]{oup-contemporary}
\documentclass[a4paper,num-refs]{oup-contemporary}
%%% Journal toggle; only specific options recognised.
%%% (Only "gigascience" and "general" are implemented now. Support for other journals is planned.)
\journal{gigascience}
\usepackage{graphicx}
\usepackage{siunitx}
\usepackage{multirow}
%%% Flushend: You can add this package to automatically balance the final page, but if things go awry (e.g. section contents appearing out-of-order or entire blocks or paragraphs are coloured), remove it!
% \usepackage{flushend}
\title{Ultra-deep, long-read nanopore sequencing of mock microbial community standards}
%%% Use the \authfn to add symbols for additional footnotes, if any. 1 is reserved for correspondence emails; then continuing with 2 etc for contributions.
\author[1\authfn{2}]{Samuel M. Nicholls}
\author[1\authfn{2}]{Joshua C. Quick}
\author[2]{Shuiquan Tang}
\author[1\authfn{1}]{Nicholas J. Loman}
\affil[1]{Institute of Microbiology and Infection, School of Biosciences, University of Birmingham, UK}
\affil[2]{Zymo Research Corporation, Irvine, California, USA}
%%% Author Notes
\authnote{\authfn{1}n.j.loman@bham.ac.uk}
\authnote{\authfn{2}Contributed equally.}
%%% Paper category
\papercat{Data Note}
%%% "Short" author for running page header
\runningauthor{Nicholls and Quick et al.}
%%% Should only be set by an editor
\jvolume{00}
\jnumber{0}
\jyear{2018}
\begin{document}
\begin{frontmatter}
\maketitle
\begin{abstract}
\textbf{Background}:
Long sequencing reads are information-rich: aiding \textit{de novo} assembly and reference mapping, and consequently have great potential for the study of microbial communities. However, the best approaches for analysis of long-read metagenomic data are unknown. Additionally, rigorous evaluation of bioinformatics tools is hindered by a lack of long-read data from validated samples with known composition. \\
\textbf{Methods}: We sequenced two commercially-available mock communities containing ten microbial species (ZymoBIOMICS Microbial Community Standards) with Oxford Nanopore GridION and PromethION. Both communities and the ten individual species isolates were also sequenced with Illumina technology. \\
\textbf{Data}: We generated 14 and 16 Gbp from two GridION flowcells and 150 and 153 Gbp from two PromethION flowcells for the evenly-distributed and log-distributed communities respectively. Read length N50 ranged between 5.3 Kbp and 5.4 Kbp over the four sequencing runs.
Basecalls and corresponding signal data are made available (4.2 TB in total).
\\
\textbf{Results}: Alignment to Illumina-sequenced isolates demonstrated the expected microbial species at anticipated abundances, with the limit of detection for the lowest abundance species below 50 cells (GridION). \textit{De novo} assembly of metagenomes recovered long contiguous sequences without the need for pre-processing techniques such as binning.
\\
\textbf{Conclusions}: We present ultra-deep, long-read nanopore datasets from a well-defined mock community. These datasets will be useful for those developing bioinformatics methods for long-read metagenomics and for the validation and comparison of current laboratory and software pipelines.
\end{abstract}
\begin{keywords}
bioinformatics; metagenomics; mock community; nanopore; single-molecule sequencing; real-time sequencing; benchmark; GridION; PromethION; Illumina; \textit{de novo} assembly
\end{keywords}
\end{frontmatter}
%%% Key points will be printed at top of second page
%\begin{keypoints*}
%\begin{itemize}
%\item This is the first point
%\item This is the second point
%\item One last point.
%\end{itemize}
%\end{keypoints*}
\section{Data Description}
Whole-genome sequencing of microbial communities (metagenomics) has revolutionised our view of microbial evolution and diversity, with numerous potential applications for microbial ecology, clinical microbiology and industrial biotechnology \cite{Handelsman2004-lg,Hug2016-wz}.
Typically, metagenomic studies use high-throughput sequencing platforms (\textit{e.g.} Illumina) \cite{Quince2017-lf} which generate very high yield, but of limited read length (100$-$300 bp).
In contrast, single-molecule sequencing platforms such as the Oxford Nanopore MinION, GridION and PromethION are able to sequence very long fragments of DNA (>10 Kbp, with over 2 Mbp reported) \cite{Jain2018-aw,Payne2018-cv} and with recent improvements to the platform making metagenomic studies using nanopore more viable, such studies are increasing in frequency \cite{Sanderson2018-zf,Charalampous2018-fl,Somerville2018-mq,Leggett2017-sx}.
Long reads help with alignment-based assignment of taxonomy and function due to their increased information content \cite{Huson2007-mk,Wommack2008-qf}.
Additionally, long reads permit bridging of repetitive sequences (within and between genomes), aiding genome completeness in \textit{de novo} assembly \cite{Bertrand2018-lz}.
However, these advantages are constrained by high error rate ($\approx$10\%), requiring the use of specific long-read alignment and assembly methods, which are either not specifically designed for metagenomics, or have not been extensively tested on real data \cite{sczyrba2017critical}.
Mock community standards are useful for the development of genomics methods \cite{Mason2017-yj}, and for the validation of existing laboratory, software and bioinformatics approaches. For example, validating the accuracy of a taxonomic identification pipeline is important, because the consequences of erroneous taxonomic identification from a metagenomic analysis may be severe, \textit{e.g.} in public health microbiology \cite{Ackelsberg2015-fh,mcintyre2017comprehensive} or incorrect diagnoses in clinical microbiology diagnostics. Mock community standards can also be used as positive controls during laboratory work, for example to validate that DNA extraction methods will yield expected representation of a sampled community \cite{Mason2017-yj}.
Here, we present four nanopore sequencing datasets of two microbial community standards, providing a state-of-the-art benchmark to accelerate the development of methods for analysing long-read metagenomics data.
\begin{table*}[t!]
\centering
\caption{Description of the ten organisms comprising the ZymoBIOMICS Mock Community Standards.}
\label{tab:strains}
\begin{tabular}{r | l | l | l | l | l | c | c c }
\toprule
& & {Est. Size} & NRRL & ATCC & Sequence & Illumina & PacBio RSII & PacBio Sequel\\
Species & Type & {(Mbp)} & Accession & Accession & Type & FASTQ & FASTQ \cite{mcintyre2019single} & FASTQ \cite{mcintyre2019single}\\
\midrule
\textit{Bacillus subtilis} & Gram $+$ & 4.045 & \texttt{B-354} & \texttt{6633}& \texttt{ST7} & \texttt{ERR2935851} & \texttt{SRR7498042} & \texttt{SRR7415629} \\
\textit{Cryptococcus neoformans} & \multirow{2}{*}{Yeast} & \multirow{2}{*}{18.9} & \multirow{2}{*}{\texttt{Y-2534}} & \multirow{2}{*}{\texttt{32045}} & \multirow{2}{*}{$-$} & \multirow{2}{*}{\texttt{ERR2935856}} & \multirow{2}{*}{$-$} & \multirow{2}{*}{$-$}\\
$\times$ \textit{Cryptococcus deneoformans} & & & & & \\
\textit{Enterococcus faecalis} & Gram $+$ & 2.845 & \texttt{B-537} & \texttt{7080}& \texttt{ST55} & \texttt{ERR2935850} & \texttt{SRR7415622} & \texttt{SRR7415630}\\
\textit{Escherichia coli} & Gram $-$ & 4.875 & \texttt{B-1109} & $-$& \texttt{ST10} & \texttt{ERR2935852} & \texttt{SRR7498041} & $-$\\
\textit{Lactobacillus fermentum} & Gram $+$ & 1.905 & \texttt{B-1840} & \texttt{14931}& $-$ & \texttt{ERR2935857} & $-$ & $-$\\
\textit{Listeria monocytogenes} & Gram $+$ & 2.992 & \texttt{B-33116} & \texttt{19117}& \texttt{ST449} & \texttt{ERR2935854} & \texttt{SRR7415624} & \texttt{SRR7415635}\\
\textit{Pseudomonas aeruginosa} & Gram $-$ & 6.792 & \texttt{B-3509} & \texttt{15442}& \texttt{ST252} & \texttt{ERR2935853} & \texttt{SRR7498043} & $-$\\
\textit{Saccharomyces cerevisiae} & Yeast & 12.1 & \texttt{Y-567} & \texttt{9763}& $-$ & \texttt{ERR2935855} & \texttt{SRR7498048} & \texttt{SRR7415638}\\
\textit{Salmonella enterica} & Gram $-$ & 4.760 & \texttt{B-4212} & $-$& \texttt{ST139} & \texttt{ERR2935848} & \texttt{SRR7415626} & \texttt{SRR7415636}\\
\textit{Staphylococcus aureus} & Gram $+$ & 2.730 & \texttt{B-41012} & $-$& \texttt{ST9} & \texttt{ERR2935849} & \texttt{SRR7415627} & \texttt{SRR7415637}\\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item Table adapted from ZymoBIOMICS™ Microbial Community Standard II (Log Distribution) Instruction Manual v1.1.2 Table 2 and Appendix A.
The \textit{S. enterica} genome is listed at NRRL (\texttt{B-4212}) as Serovar Typhimurium LT2, but our genomic analysis shows it is likely to be Serotype Choleraesuis; indicating possible mis-annotation.
%PacBio RSII and Sequel sequencing were conducted as part of another study by McIntyre \textit{et al.} \cite{mcintyre2019single}.
\end{tablenotes}
\end{table*}
\begin{table*}[b!]
\centering
\caption{Summary of the four nanopore sequencing experiments.}\label{tab:datasets}
\begin{tabular}{l l | l l | c c S S | S S}
\toprule
Signal & FASTQ & & & {Time} & {Reads} & {N50} & {Quality} & {Yield} & {Q$>$7} \\
Accession&Accession & Sequencer & Standard (Lot) & {(h)} & {(M)} & {(Kbp)} & {(Median Q)} & {(Gbp)} & {(Gbp)} \\
\midrule
\texttt{ERR2887847} & \texttt{ERR3152364} & GridION & Zymo CS Even \texttt{ZRC190633} & 48 & 3.49 & 5.3 & 10.3 & 14.38 & 12.39 \\
\midrule
\texttt{ERR2887850} &\texttt{ERR3152366} & GridION & Zymo CSII Log \texttt{ZRC190842} & 48 & 3.67 & 5.4 & 9.8 & 16.51 & 13.97 \\
\midrule
\texttt{ERR2887848} & \multirow{2}{*}{\texttt{ERR3152365}}& PromethION& Zymo CS Even \texttt{ZRC190633} & 64 &{\multirow{2}{*}{35.7}} & {\multirow{2}{*}{5.4}} & {\multirow{2}{*}{10.5}} & {\multirow{2}{*}{150.88}} & {\multirow{2}{*}{130.32}}\\
\texttt{ERR2887849} & & PromethION& Zymo CS Even \texttt{ZRC190633} & -- & & & & \\
\midrule
\texttt{ERR2887851} & \multirow{2}{*}{\texttt{ERR3152367}} & PromethION& Zymo CSII Log \texttt{ZRC190842} &64& {\multirow{2}{*}{34.5}} & {\multirow{2}{*}{5.4}} & {\multirow{2}{*}{10.7}} & {\multirow{2}{*}{153.31}} & {\multirow{2}{*}{133.68}} \\
\texttt{ERR2887852} && PromethION& Zymo CSII Log \texttt{ZRC190842} &--& & & & \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item PromethION runs were restarted following the standard 64 hour protocol. The table reflects total yield across both the standard run and subsequent restarts.
\end{tablenotes}
\end{table*}
\subsection{Background Information}
The ZymoBIOMICS Microbial Community Standards (CS and CSII) are each composed of ten microbial species: eight bacteria and two yeasts (Table \ref{tab:strains}). The organisms in CS (hereafter referred to as `Even') are distributed equally (12\%), with the exception of the two yeasts which are each present at 2\%.
Cell counts from organisms in CSII (`Log') community are distributed on a log scale, ranging from 89.1\% (\textit{Listeria monocytogenes}), down to 0.000089\% (\textit{Staphylococcus aureus}).
%The standards are certified by the manufacturer to contain no more than 0.01\% contaminating cells.
\section{Methods}
\subsection{DNA extraction}
DNA was extracted from \SI{75}{\micro\litre} ZymoBIOMICS Microbial Community Standard (Product D6300, Lot ZRC190633) and \SI{375}{\micro\litre} ZymoBIOMICS Microbial Community Standard II (Product D6310, Lot ZRC190842) using the ZymoBIOMICS DNA Miniprep extraction kit according to manufacturer's instructions, with the following modifications to increase fragment length and maintain the expected representation of the Gram-negative species which are already lysed in the DNA/RNA Shield storage solution.
The standard was centrifuged at 8,000$\times$\textit{g} for 5 minutes before removing the supernatant and retaining. The cell pellet was resuspended in \SI{750}{\micro\litre} lysis buffer and added to the ZR BashingBead lysis tube. Bead-beating was performed on a FastPrep-24 (MP Biomedicals) instrument for 2 cycles of 40 seconds at \SI{6.0}{\metre\per\second}, with 5 minutes sitting on ice between cycles. The bead tubes were centrifuged at 10,000$\times$\textit{g} for 1 minute and \SI{450}{\micro\litre} of supernatant was transferred to a Zymo Spin III-F filter before being centrifuged again at 8000$\times$\textit{g} for 1 minute.
\SI{45}{\micro\litre} (Even) and \SI{225}{\micro\litre} (Log) of the supernatant retained earlier was combined with \SI{450}{\micro\litre} filtrate before adding \SI{1485}{\micro\litre} (Even) or \SI{2025}{\micro\litre} (Log) Binding Buffer and mixing before loading onto the column.
Methods are available via \url{dx.doi.org/10.17504/protocols.io.x9tfr6n}.
\subsection{Nanopore sequencing library preparation}
Quantification steps were performed using the dsDNA HS assay for Qubit.
DNA was size-selected by cleaning up with 0.45$\times$ volume of Ampure XP (Beckman Coulter) and eluted in \SI{100}{\micro\litre} EB (Qiagen). Libraries were prepared from \SI{1400}{\nano\gram} input DNA using the SQK-LSK109 kit (Oxford Nanopore Technologies) as per manufacturer's protocol, except incubation times for end-repair, dA-tailing and ligation were increased to 30 minutes to improve ligation efficiency. The even and log libraries were split and used on both the GridION and PromethION flowcells.
\subsection{Sequencing}
Sequencing libraries were quantified and two aliquots of \SI{50}{\nano\gram} and \SI{400}{\nano\gram} were prepared for GridION and PromethION sequencing respectively. The GridION sequencing was performed using \texttt{FLO-MIN106} (rev.C) flowcells, \texttt{MinKNOW} 1.15.1 and standard 48-hour run script with active channel selection enabled. The PromethION sequencing was performed using \texttt{FLO-PRO002} flowcells, \texttt{MinKNOW} 1.14.2 and standard 64-hour run script with active channel selection enabled.
Refuelling was performed approximately every 24h (GridION, PromethION) by loading \SI{75}{\micro\litre} (GridION) or \SI{150}{\micro\litre} (PromethION) refuelling mix (SQB diluted 1:1 with nuclease-free water). Additionally, after the standard scripts had completed the PromethION was restarted several times to utilise remaining active pores and maximise total yield.
\subsection{Nanopore basecalling}
Reads were basecalled on-instrument using the \texttt{Guppy} v2.2.2 GPU basecaller (Oxford Nanopore Technologies) with the supplied \texttt{dna\_r9.4.1\_450bps\_flipflop\_prom.cfg} configuration (PromethION) and \texttt{dna\_r9.4.1\_450bps\_flipflop.cfg} (GridION).
\subsection{Illumina sequencing}
DNA was extracted from pure cultures of each species using the ZymoBIOMICS DNA Miniprep Kit. Library preparation was performed using the Kapa HyperPlus Kit with \SI{100}{\nano\gram} DNA as input and TruSeq Y-adapters. The purified library derived from each sample was quantified by TapeStation (Agilent 4200) and pooled together in an equimolar fashion. The multiplexed isolates were sequenced on an Illumina HiSeq 1500 instrument using 2$\times$101 bp (paired-end) sequencing, over four lanes. Raw reads were demultiplexed using \texttt{bcl2fastq} v2.17.
Shotgun sequencing of the even and log communities was performed with the same protocol, with the exception that the log community was sequenced individually on two flowcell lanes, and the even community was instead sequenced on an Illumina MiSeq using 2$\times$151 bp (paired-end) sequencing.
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\begin{figure*}[t!]
\centering
\includegraphics[width=0.975\linewidth]{figures/Qual.png}
\caption{
Summary plots for the four generated data sets:
(a) collector's curve showing sequencing yield over time for each of the four sequencing runs,
(b) density plot showing sequence accuracy (\texttt{BLAST}-like identities),
(c) density plot showing sequencing speed over time by sequencing experiment.
}\label{fig:summaries}
\end{figure*}
\begin{figure}[b!]
\centering
\includegraphics[width=0.905\linewidth]{figures/log-ove.png}
\caption{Proportion of sequenced bases assigned by \texttt{minimap2} to each of the 10 organisms that were sequenced (x-axis), against the proportion of yield expected given the known composition (y-axis) of the Zymo CSII (Log) standard.}\label{fig:log-ove}
\end{figure}
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\begin{table}[b!]
\centering
\caption{Summary statistics for Illumina sequencing data.}\label{tab:illumina}
\begin{tabular}{l | S S S c}
\toprule
{} & {Pairs} & {Yield} & {{phred}$\geq$30} & {}\\
{Dataset} & {(M)} & {(Gbp)} & {(\%)} & {Accession}\\
\midrule
\multirow{2}{*}{Isolates} & 13.53 & 2.73 & 87.72\% & \multirow{2}{*}{see Table \ref{tab:strains}}\\
& {$\pm$} 5.23 & {$\pm$} 1.06 & {$\pm$} 5.43\% & \\
\midrule
{CS (Even)} & 8.8 &2.65 & 95.12\% & \texttt{ERR2984773}\\
{CSII (Log)} & 47.8 &9.66 & 95.71\% & \texttt{ERR2935805}\\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item Illumina sequencing was performed on an Illumina HiSeq 1500, with the exception of the even community which was sequenced on an Illumina MiSeq.
\end{tablenotes}
\end{table}
\section{Bioinformatics Methods}
\subsection{Illumina draft assembly}
For the purposes of estimating sequencing coverage and contiguity, we constructed a draft assembly from our available Illumina sequencing data.
Illumina reads for each of the ten isolates were assembled using \texttt{SPAdes} v3.12.0 \cite{Bankevich2012-iu} with paired-end reads as input, using parameters \texttt{-m 512 -t 12}.
Scaffolds from \texttt{SPAdes} less than 500 bp length or with less than 10$\times$ coverage were removed. The remaining scaffolds were combined into a single mock community draft assembly for downstream analysis.
Multilocus sequence typing (MLST) of the scaffolds was conducted with \texttt{mlst} (\url{https://github.com/tseemann/mlst}).
\subsection{PacBio draft assembly}
A recently released orthogonal data set from McIntyre \textit{et al}.\ (2019) includes individual PacBio sequencing of eight of the ten organisms that compose the two Zymo communities \cite{mcintyre2019single}.
Assemblies for the eight isolates that passed quality control (excluding \textit{L. fermentum} and \textit{C. neoformans}) were generated with \texttt{HGAP v2} \cite{chin2013nonhybrid}.
Assemblies have been made available by the authors and were downloaded from {\url{https://github.com/al-mcintyre/mCaller_analysis_scripts/tree/master/assemblies}} (\texttt{Git} commit \texttt{dba494d}) for the purposes of assessing metagenomic assembly accuracy for the 7 bacterial species where complete genomes were available.
\subsection{Sequencing coverage estimation}
Nanopore reads were aligned to the Illumina draft assembly using \texttt{minimap2} \cite{li2018minimap2} v2.14-r883 with parameters \texttt{-ax map-ont -t 12} and converted to a sorted BAM file using \texttt{samtools} \cite{li2009sequence}.
To reduce erroneous mappings, alignment BAM files were filtered using a script \texttt{bamstats.py} according to the following criteria; reference mapping length $\geq$500 bp, map quality (MAPQ) > 0, there are no supplementary alignments for this read and read is not a secondary alignment.
Per-species coverage summary statistics were generated using the \texttt{summariseStats.R} Rscript.
\begin{table*}[t!]
\caption{Read alignment statistics for even samples, showing absolute measurements and proportion of sequencing yield and the estimated genome coverage obtained for each organism in the mock community.}\label{tab:mappings-even}
\begin{tabularx}{\linewidth}{l c | S S c S | S S c S }
\toprule
{} & {} & \multicolumn{4}{c|}{GridION} & \multicolumn{4}{c}{PromethION} \\
{} & {Expected} & {Yield} & {Measured} & {Aln. N50} & {Coverage} & {Yield} & {Measured} & {Aln. N50} & {Coverage} \\
{Species} & {Proportion} & {(Gbp)} & {Proportion} & {(Kbp)} & {($\times$)} & {(Gbp)} & {Proportion} & {(Kbp)} & {($\times$)} \\
\midrule
\textit{Bacillus subtilis} &12& 2.12 &19.32 &4.30& 524.51 & 21.55 &19.02 &4.40& 5326.44 \\
\textit{Listeria monocytogenes} &12& 1.60 &14.56 &4.47& 534.26 & 16.23 &14.33 &4.58& 5424.46 \\
\textit{Enterococcus faecalis} &12& 1.34 &12.24 &4.45& 472.47 & 13.67 &12.07 &4.57& 4805.60 \\
\textit{Staphylococcus aureus} &12& 1.24 &11.28 &4.47& 453.84 & 12.59 &11.11 &4.59& 4611.61 \\
\textit{Salmonella enterica} &12& 1.10 &9.99 &8.55& 230.51 & 11.69 &10.32 &8.95& 2456.19 \\
\textit{Escherichia coli} &12& 1.09 &9.93 &8.31& 223.59 & 11.62 &10.26 &8.71& 2382.59 \\
\textit{Pseudomonas aeruginosa} &12& 1.07 &9.70 &8.98& 156.85 & 11.45 &10.11 &9.38& 1686.34 \\
\textit{Lactobacillus fermentum} &12& 1.02 &9.28 &3.62& 534.73 & 10.34 &9.13 &3.73& 5425.69 \\
\textit{Saccharomyces cerevisiae} &2& 0.21 &1.92 &4.09& 17.46 & 2.12 &1.87 &4.18& 175.23 \\
\textit{Cryptococcus neoformans} &2& 0.20 &1.78 &4.45& 10.37 & 2.00 &1.77 &4.54& 105.82 \\
\bottomrule
\end{tabularx}
%\begin{tablenotes}
%\item
%\end{tablenotes}
\end{table*}
\begin{table*}[tb!]
\centering
\caption{Read alignment statistics for log samples, describing sequencing yield and estimated genome coverage obtained for each organism in the mock community.}\label{tab:mappings-odd}
\begin{tabular}{l | S c S | S c S }
\toprule
{} & \multicolumn{3}{c|}{GridION} & \multicolumn{3}{c}{PromethION} \\
{} & {Yield} & {Aln. N50} & {Coverage} & {Yield} & {Aln. N50} & {Coverage} \\
{Species} & {(Gbp)} & {(Kbp)} & {($\times$)} & {(Gbp)} & {(Kbp)} & {($\times$)} \\
\midrule
\textit{Listeria monocytogenes} & 12.10 &4.95& 4043.90 & 110.09 &4.97& 36796.21\\
\textit{Pseudomonas aeruginosa} & 1.10 &9.38& 161.45 & 9.99 &9.33& 1471.41\\
\textit{Bacillus subtilis} & 0.16 &5.03& 38.67 & 1.44 &5.04& 356.00 \\
\textit{Saccharomyces cerevisiae} & 0.08 &4.78& 6.93 & 0.75 &4.75& 62.33\\
\textit{Salmonella enterica} & 0.01 &9.20& 2.20 & 0.10 &9.17& 20.04 \\
\textit{Escherichia coli} & 0.01 &8.65& 2.14 & 0.09 &9.17& 19.24 \\
\textit{Lactobacillus fermentum} & 4E-4 &3.40& 0.210 & 0.004 &3.37& 2.03 \\
\textit{Enterococcus faecalis} & 2E-4 &7.62& 0.055 & 1E-3 &6.05& 0.34 \\
\textit{Cryptococcus neoformans} & 6E-5 &4.41& 0.003 & 7E-4 &4.97& 0.037\\
\textit{Staphylococcus aureus} & 1E-5 &7.12& 0.005 & 5E-5 &3.58& 0.020 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item Note that expected and measured proportions are illustrated by Figure \ref{fig:log-ove}.
\end{tablenotes}
\end{table*}
\subsection{Nanopore read accuracy}
Read accuracy was determined by calculating \texttt{BLAST}-like identities from the filtered alignments (as per \url{http://lh3.github.io/2018/11/25/on-the-definition-of-sequence-identity}), calculated as $(L-NM)/L$ using the \texttt{minimap2} number of mismatches (\texttt{NM}) SAM tag and the sum of match, insertion and deletion CIGAR operations (\texttt{L}).
\subsection{Metagenomic assembly and contiguity estimation}
Metagenomic assemblies were constructed with \texttt{wtdbg2} v2.2 \cite{ruan2019fast} from the nanopore sequencing of the communities.
\texttt{wtdbg2} was compiled from source via \texttt{Git} commit \texttt{904f2b3}. For GridION, all nanopore reads were used. For PromethION, a 25\% subsample was selected with \texttt{seqtk} (\url{https://github.com/lh3/seqtk}).
Assemblies were conducted under a variety of parameter values for homopolymer-compressed k-mer size (\texttt{-p}), minimum graph edge weight support (\texttt{-e}) and read length threshold (\texttt{-L}). Global parameters for all runs (\texttt{-S1 -K10000 -{}-node-max 6000}) were used to turn-off k-mer subsampling (to remove assembly stochasticity) and increase the coverage thresholds applied to k-mers and constructed nodes.
Assembled contigs were assigned to taxa with \texttt{kraken2} \cite{Wood2014} (\texttt{-{}-use-names -t12}) using a database containing all of the archaeal, bacterial, fungal, protozoal and viral sequences from \texttt{RefSeq}, and \texttt{UniVec\_Core} (database download links are in our repository).
The \texttt{kraken2} output was parsed with \texttt{extracken.py} and plotted with \texttt{contiguity.R} to visually assess contiguity.
Following assignment, contigs can be extracted into separate FASTA with \texttt{extract\_contigs\_with\_kraken.py}.
\subsection{Assembly polishing}
After inspection of the \texttt{contiguity.R} plot, eight high-contiguity assemblies were selected for polishing.
Polishing consisted of two iterations of \texttt{racon} \cite{vaser2017fast}, followed by \texttt{medaka} (\url{https://github.com/nanoporetech/medaka}) and two iterations of \texttt{pilon} \cite{walker2014pilon}.
\texttt{racon} v1.3.2 was used to polish contigs with the the nanopore reads.
\texttt{medaka} v0.5.0 was used to polish the \texttt{racon} polished contigs, with the nanopore reads specifying the \texttt{r941\_flip} model.
The PromethION assemblies were polished using the same \texttt{seqtk}-derived 25\% subset from which the assemblies were constructed.
\texttt{pilon} v1.23 was used to polish the \texttt{medaka} polished contigs, with the CS (Even) community Illumina reads.
\subsection{Estimation of genome completeness}
To estimate accuracy of the polished assemblies, contigs were first assigned to taxa and extracted into separate FASTA using \texttt{kraken2} as previously described.
For the seven bacteria for which corresponding PacBio draft assemblies were available, sequence identity dotplots were generated using a modified version of \texttt{minidot} (\url{https://github.com/SamStudio8/minidot}) which uses \texttt{minimap2} (\texttt{-x asm10 -{}-no-long-join -{}-dual=yes -P}) to align the polished contigs binned by \texttt{kraken2}, to the corresponding PacBio draft.
Genome completeness was estimated with \texttt{CheckM} v1.0.13 \cite{parks2015checkm} using the \texttt{taxonomy\_wf} subcommand, after each phase of the polishing pipeline.
\texttt{CheckM} was executed separately for each \texttt{kraken2} bin that had a corresponding PacBio reference, specifying the appropriate species for the bin to \texttt{taxonomy\_wf}.
We report the \texttt{CheckM} ``Completeness'' score, which estimates completeness by identifying collocated marker gene sets on the assembled contigs as a proportion of the total
collection of marker gene sets expected for a specific taxon.
\begin{figure*}[t!]
\centering
\includegraphics[width=\linewidth]{figures/w2.png}
\caption{Bar plots demonstrating total length and contiguity of genomic assemblies obtained with \texttt{wtdbg2} from each of the long-read nanopore data sets. For each organism in the community (coloured columns), contigs longer than 10 kbp are horizontally stacked along the x-axis. Each row represents a run of \texttt{wtdbg2}, with the parameters for edge support, read length threshold and homopolymer-compressed k-mer size labelled on the left. Assemblies are grouped by the data set on which they were run (row facets). Additionally, assemblies may be compared to the estimated true genome size, the available McIntyre \textit{et al}.\ PacBio assemblies, and per-isolate Illumina \texttt{SPAdes} assembly. Estimated genomes sizes are the same as those found in Table \ref{tab:strains}, however to display approximate chromosomes, the two yeasts were replaced by their corresponding canonical NCBI references for visualisation purposes only. The \textit{C. neoformans} strain used by the Zymo standards is a diploid genetic cross, which may explain the larger assemblies, compared to the represented estimated haploid size.
}\label{fig:assemblies}
\end{figure*}
\section{Results}
\subsection{Nanopore sequencing metrics}
We generated a total of 335.1 Gbp of sequence from the four nanopore sequencing runs (Table \ref{tab:datasets}, Figure \ref{fig:summaries}a).
PromethION flowcells generated approximately ten-times more sequencing data than the comparative GridION runs and showed equivalent read length N50 and read accuracy (Figure \ref{fig:summaries}b).
We observe a difference in sequencing speed between the PromethION (mean 419 bps and 437 bps for even and log) and the GridION (mean speed 352 and 372 bps for even and log) (Figure \ref{fig:summaries}c).
\subsection{Illumina sequencing metrics}
Illumina datasets for the ten individually sequenced isolates averaged 13.53 million pairs of reads (ranging between 7.1 $-$ 23.2 million), with proportions of reads with a mean \texttt{phred} score $\geq$30 ranging between 75.51\% $-$ 93.09\% (Table \ref{tab:illumina}).
Illumina sequencing generated 8.8 million pairs of reads (2$\times$151 bp, MiSeq), and 47.8 million pairs of reads (2$\times$101 bp, HiSeq) for the even and log community, respectively (Table \ref{tab:illumina}).
\setlength{\tabcolsep}{0.5pt}
\begin{table*}[t!]
\caption{Sequence identity dotplots and \texttt{CheckM} genome completeness scores for each of the seven bacteria for which there was a corresponding PacBio assembly from McIntyre \textit{et al}.\ (2019). Four \texttt{wtdbg2} assembly conditions are represented, varying the homopolymer-compressed k-mer parameter `\texttt{p}' and the graph minimum edge weight threshold `\texttt{e}'. The read length threshold `\texttt{L}' was fixed at 5000 bp. The left and right halves of the table correspond to the same assembly condition for the GridION and 25\% PromethION sequencing data, respectively. The L50/L95 refers to the number of assembled contigs required to span at least 50\% and 95\% of the estimated genome size (see Table \ref{tab:strains}).\\An $-$ indicates that the set of assembled contigs assigned to a taxon were not of sufficient total length to cover 95\% of the estimated size. \texttt{CheckM} genome completeness scores are expressed as a percentage and were calculated per-organism at the end of each polishing phase.}
\label{tab:dotplot}
\begin{tabular}{ccccccc|cc|ccccccc}
\multicolumn{7}{c}{GridION} & \multicolumn{2}{c}{} & \multicolumn{7}{c}{PromethION} \\
bs & ef & ec & lm & pa & se & sa & \multicolumn{2}{c|}{Assembly} & bs & ef & ec & lm & pa & se & sa\\
\toprule
\multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{0d0198ab-13a0-4758-8483-2b51d147dbed.ctg.cns.staphylococcus_aureus}.png}} & \ \ -p & 21 & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{58d40d60-8fee-4b2b-96cc-a12beaad77db.ctg.cns.staphylococcus_aureus}.png}}\\
& & & & & & & \ \ -e & 3 & & & & & & & \\
& & & & & & & & & & & & & & & \\
2 / 5 & 1 / 1 & 5 / 17 & 1 / 3 & 1 / 1 & 6 / 21 & 1 / 1 & \multicolumn{2}{c|}{L50/L95} & 2 / 5 & 1 / 2 & 4 / $-$ & 3 / 8 & 2 / 5 & 4 / 15 & 1 / 2\\
74.27 & 76.10 & 70.14 & 76.11 & 78.64 & 66.78 & 78.48 & \multicolumn{2}{c|}{Base} & 70.25 & 72.82 & 62.02 & 69.46 & 78.17 & 65.01 & 74.35\\
86.65 & 88.08 & 84.21 & 83.41 & 92.96 & 82.74 & 90.27 & \multicolumn{2}{c|}{+Racon$\times$2} & 83.20 & 83.93 & 71.02 & 82.33 & 90.38 & 77.68 & 86.97\\
97.45 & 99.07 & 94.46 & 97.50 & 97.33 & 95.10 & 98.35 & \multicolumn{2}{c|}{+Medaka} & 95.83 & 97.74 & 81.65 & 96.70 & 99.14 & 91.48 & 97.69\\
98.42 & 99.66 & 95.46 & 98.57 & 99.77 & 96.98 & 98.88 & \multicolumn{2}{c|}{+Pilon$\times$2} & 98.44 & 99.66 & 83.59 & 98.66 & 99.73 & 95.02 & 98.20\\
\midrule
\multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{7cd60d3b-eafb-48d1-9aab-c8701232f2f8.ctg.cns.staphylococcus_aureus}.png}} & \ \ -p & 23 & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{a521bcba-06e5-4e48-b5da-c7808a6588e8.ctg.cns.staphylococcus_aureus}.png}}\\
& & & & & & & \ \ -e & 3 & & & & & & & \\
& & & & & & & & & & & & & & & \\
2 / 4 & 1 / 1 & 2 / $-$ & 1 / 2 & 1 / 1 & 2 / 4 & 1 / 1 & \multicolumn{2}{c|}{L50/L95} & 2 / 3 & 1 / 2 & 2 / $-$ & 1 / 3 & 1 / 1 & 3 / 7 & 1 / 1\\
73.42 & 75.81 & 67.44 & 74.44 & 81.08 & 70.54 & 78.20 & \multicolumn{2}{c|}{Base} & 68.20 & 71.32 & 59.15 & 70.02 & 79.85 & 65.62 & 74.04\\
84.81 & 87.30 & 79.83 & 84.57 & 92.53 & 85.83 & 88.64 & \multicolumn{2}{c|}{+Racon$\times$2} & 83.23 & 83.36 & 67.16 & 82.42 & 87.72 & 77.46 & 84.17\\
96.60 & 98.80 & 89.41 & 98.06 & 98.04 & 97.46 & 98.23 & \multicolumn{2}{c|}{+Medaka} & 95.98 & 97.71 & 79.47 & 97.42 & 99.03 & 94.14 & 97.90\\
97.83 & 99.66 & 90.36 & 99.15 & 99.82 & 98.67 & 98.88 & \multicolumn{2}{c|}{+Pilon$\times$2} & 98.34 & 99.66 & 81.65 & 99.27 & 99.77 & 98.18 & 98.88\\
\midrule
\multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{2c59020c-85f2-4bf3-bfaf-314f8fc41cdc.ctg.cns.staphylococcus_aureus}.png}} & \ \ -p & 21 & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{71740786-24eb-4a16-aa27-c9bdf9dce3e2.ctg.cns.staphylococcus_aureus}.png}}\\
& & & & & & & \ \ -e & 10 & & & & & & & \\
& & & & & & & & & & & & & & & \\
1 / 1 & 1 / 2 & 3 / 8 & 1 / $-$ & 1 / 1 & 2 / 7 & 1 / 1 & \multicolumn{2}{c|}{L50/L95} & 2 / 5 & 2 / 4 & 3 / 14 & 2 / 7 & 1 / 2 & 3 / 10 & 1 / 2\\
74.29 & 76.07 & 72.21 & 57.20 & 79.64 & 68.99 & 78.28 & \multicolumn{2}{c|}{Base} & 70.82 & 72.29 & 67.56 & 71.91 & 79.04 & 66.33 & 74.20\\
85.36 & 86.11 & 84.53 & 62.94 & 92.28 & 84.46 & 90.46 & \multicolumn{2}{c|}{+Racon$\times$2} & 83.88 & 85.62 & 77.20 & 84.26 & 90.35 & 79.43 & 88.35\\
97.14 & 99.11 & 95.95 & 71.21 & 98.06 & 96.57 & 98.55 & \multicolumn{2}{c|}{+Medaka} & 96.87 & 97.57 & 90.88 & 97.43 & 98.92 & 95.51 & 98.01\\
98.27 & 99.66 & 97.17 & 72.24 & 99.53 & 98.34 & 98.78 & \multicolumn{2}{c|}{+Pilon$\times$2} & 98.43 & 99.66 & 92.58 & 99.15 & 99.72 & 97.85 & 98.86\\
\midrule
\multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{c8e73aed-b0cb-40dc-b9ff-654a6b7c7805.ctg.cns.staphylococcus_aureus}.png}} & \ \ -p & 23 & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.bacillus_subtilis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.enterococcus_faecalis}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.escherichia_coli}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.listeria_monocytogenes}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.pseudomonas_aeruginosa}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.salmonella_enterica}.png}} & \multirow{3}{*}[0.05cm]{\includegraphics[width=1.1cm]{figures/minidotplots/{98840cca-6e42-4167-8d55-5b92835dbb34.ctg.cns.staphylococcus_aureus}.png}}\\
& & & & & & & \ \ -e & 10 & & & & & & & \\
& & & & & & & & & & & & & & & \\
2 / 3 & 1 / 1 & 2 / 3 & 1 / 4 & 1 / 2 & 2 / $-$ & 1 / 3 & \multicolumn{2}{c|}{L50/L95} & 1 / 3 & 2 / 4 & 1 / 4 & 1 / 4 & 1 / 1 & 1 / 4 & 2 / 3\\
73.71 & 77.52 & 73.73 & 75.42 & 82.06 & 60.74 & 79.24 & \multicolumn{2}{c|}{Base} & 69.27 & 71.85 & 70.35 & 71.17 & 80.41 & 67.06 & 74.98\\
86.42 & 88.31 & 84.69 & 85.14 & 92.81 & 71.29 & 87.81 & \multicolumn{2}{c|}{+Racon$\times$2} & 81.55 & 84.89 & 82.92 & 83.81 & 89.17 & 80.33 & 88.61\\
97.16 & 98.83 & 94.26 & 96.86 & 97.86 & 82.42 & 98.45 & \multicolumn{2}{c|}{+Medaka} & 96.62 & 98.46 & 95.42 & 96.40 & 98.72 & 96.26 & 97.74\\
98.44 & 99.66 & 98.13 & 98.69 & 99.83 & 83.58 & 98.86 & \multicolumn{2}{c|}{+Pilon$\times$2} & 98.33 & 99.66 & 97.14 & 98.69 & 99.72 & 98.75 & 98.78\\
\midrule
bs & ef & ec & lm & pa & se & sa & \multicolumn{2}{c|}{} & bs & ef & ec & lm & pa & se & sa\\
\end{tabular}
\end{table*}
\subsection{Nanopore mapping statistics}
We identify the presence of all 10 microbial species in the community, for both even and log samples, in expected proportions (Figure \ref{fig:log-ove}). For the even community, the GridION results provide sufficient depth (i.e. $\gg30\times$ coverage) to potentially assemble all eight of the bacteria. The coverage of the yeast genomes were lower (10$\times$ and 17$\times$), potentially sufficient for assembly scaffolding.
On the PromethION all genomes had >100$\times$ mean coverage (Tables \ref{tab:mappings-even} and \ref{tab:mappings-odd}).
For the log-distributed community, three taxa have sufficient coverage for assembly on GridION, compared with four on PromethION. On PromethION, a further two genomes (\textit{S. enterica} and \textit{E. coli}) have sufficient coverage for assembly scaffolding. We are able to detect \textit{S. aureus}, the lowest abundance organism on both platforms, with 19 reads from PromethION (from 400 cell input) and 4 reads from GridION (from 50 cell input).
\subsection{Nanopore metagenomic assemblies}
We assessed the contiguity of our nanopore metagenomic assemblies for each run with different assembly parameters.
For the even community, genomes of the expected size were present for each of the bacterial species, contained in small numbers of large contigs (Figure \ref{fig:assemblies}). However, the two yeasts are highly fragmented, consistent with their low read depth.
\textit{L. monocytogenes} is poorly assembled in the log dataset despite being the most abundant organism, indicating very high sequence coverage may be detrimental to the performance of \texttt{wtdbg2}. We note that assembling the entire PromethION dataset resulted in less complete and more fragmented assemblies.
This led us to random subsample the PromethION data to 25\% of the total dataset which improved the assembly results.
%\vfill\null
%\noindent
After subsampling, assemblies of the even community from the GridION and PromethION are similar. However, the assemblies from PromethION data had better representation of the yeasts in terms of size and contiguity (particularly for \textit{C. neoformans}), likely due to the higher coverage of these species.
We also assessed the completeness of polished genomes for a selection of our highly-contiguous metagenomic assemblies.
For the GridION, we observe for at least one of the polished assemblies, four bacterial genomes are reconstructed to at least 95\% of their length (L95) in a single contig. For PromethION, we observe that for seven bacteria, at least half the genome (L50) is reconstructed on a single contig, for at least one assembly condition (Table \ref{tab:dotplot}).
Genome completeness as estimated by \texttt{CheckM} averaged 73.95\% and 70.98\% over the four unpolished assemblies, for the GridION and PromethION respectively.
We observed each phase of the polishing pipeline improved completeness.
For the GridION assemblies, completeness was incrementally improved by 11.57 pp, 10.14 pp and 1.25 pp for two iterations of \texttt{racon}, one iteration of \texttt{medaka} and two iterations of short read polishing with \texttt{pilon}, respectively.
For the PromethION, the three polishing phases incrementally improved assemblies by an average of 11.92 pp, 12.69 pp and 1.77 pp.
In almost all cases, polishing yields near complete genomes ($\geq$90\%) genomes.
%\vfill\null
%\pagebreak
\section{Discussion}
There are several noteworthy aspects of this dataset: We generated over 300 Gbp of sequence data from the Oxford Nanopore PromethION and 30 gigabases from the Oxford Nanopore GridION, on a well-characterised mock community sample and we have made basecalls and electrical signal data for each of the four runs presented here available: a combined dataset size of over four terabytes. The availability of the raw signal permits future basecalling of the data (an area under rapid development), as well as signal-level polishing and the detection of methylated bases \cite{Simpson2017-gn}.
Individual sequencing libraries were split between the GridION and PromethION, permitting direct comparisons of the instruments to be made. We observed high concordance between the datasets from each platform. We note the sequencing speed of the PromethION is faster than the GridION, which we attribute to different running temperatures on these instruments (39$^{\circ}$C versus 34$^{\circ}$C, respectively).
Confident detection of \textit{S. aureus} was demonstrated for the GridION run to <50-cells using the log community.
The PromethION generated around five times more \textit{S. aureus} reads as the GridION, however we loaded eight times as much library, making it appear less sensitive.
It may be possible to reduce the input to PromethION flowcells, but we have not attempted this.
Early results of metagenomic assembly show promise for reconstruction of whole microbial genomes from mixed samples without a binning step. We focused on the developing \texttt{wtdbg2} software as the established \texttt{minimap2} and \texttt{miniasm} method resulted in excessively large intermediate files (tens of terabases per analysis) which were impractical to store and analyse.
For the even community, using \texttt{wtdgb2} with varying parameter choices, we were able to assemble four of the bacteria into single contigs. However, no single parameter set was found to be optimum for both total genome size and contig length.
Increasing \texttt{-e} improved contiguity for the even community, however this resulted in the loss of yeasts from the assembly. Increasing the read length threshold (\texttt{-L}) improved contiguity for all sample and platform combinations, at the cost of genome size. Increasing the homopolymer-compressed k-mer size (\texttt{-p}) from the default of 21 to 23 also appears to improve contiguity.
We found that \texttt{wtdbg2} expects a maximum of 200$\times$ sample coverage, and discards sequence k-mers and \textit{de Bruijn} graph nodes with more than 200$\times$ support. Although these limits can be lifted by specifying higher \texttt{-K} and \texttt{--node-max}, we still observe more fragmented assemblies on the PromethION data (especially for the 100\% PromethION data [not shown]) potentially indicating a need to further tune the algorithm to account for the large differences in coverage between genomes.
It should be noted that \texttt{wtdbg2} is still under active development, making it difficult to make concrete recommendations for parameters.
We found that any form of polishing improves the completeness of assemblies, likely due to the correction of frameshifts caused by indels.
Short read polishing with \texttt{pilon} also improves the assemblies, despite low coverage of the Illumina even community data and the results might be expected to improve further with increased coverage.
The availability of this dataset should help with further improvements to long-read assembly techniques.
%This study has several shortcomings:
%We have not yet explored polishing techniques to improve consensus accuracy of assemblies using nanopore or Illumina data \cite{Rang2018-md}. A new ``flip-flop'' basecaller, \texttt{Flappie} (\url{https://github.com/nanoporetech/flappie}), was recently made available although we have not used it on this dataset.
%Although reference genomes for the ten microbial strains contained in the standards are available from Zymo (\url{https://s3.amazonaws.com/zymo-files/BioPool/ZymoBIOMICS.STD.refseq.v2.zip}), they are constructed from a combination of nanopore and Illumina data (not presented here). To avoid a circular comparison, we chose to perform analysis only against Illumina draft genomes.
Other mock microbial samples are available which we did not test here. A notable alternative mock community sample is from the Human Microbiome Project (HMP) and consists of 20 microbial samples (available from BEI Resources). This mock community have been sequenced as part of other studies, although the datasets are much smaller than the ones presented here \cite{Leggett2017-sx,Huson2018-bh}. Bertrand \textit{et al}.\ presented a synthetic mock community of their own construction to demonstrate hybrid nanopore-Illumina metagenome assemblies \cite{Bertrand2018-lz}.
\subsection{Re-use potential}
The provision of Illumina reads for each isolate permits a ground-truth to be obtained for the individual species contained in the mock community. This will be useful for training new nanopore basecalling and polishing models, long-read aligners, variant callers, and validating taxonomic assignment and assembly software and pipelines.
\section{Availability of source code and requirements}
Python and R scripts used to generate the summary information and analyses are open source and freely available via our repository (\url{https://github.com/LomanLab/mockcommunity}), under the MIT license.
Our pipeline was orchestrated with \texttt{Snakemake} \cite{koster2012snakemake}, the workflow is available from our repository.
%Lists the following:
%\begin{itemize}
%\item Project name: e.g.~My bioinformatics project
%\item Project home page: e.g.~\url{http://sourceforge.net/projects/mged}
%\item Operating system(s): e.g.~Platform independent
%\item Programming language: e.g.~Java
%\item Other requirements: e.g.~Java 1.3.1 or higher, Tomcat 4.0 or higher
%\item License: e.g.~GNU GPL, FreeBSD etc.
%Any restrictions to use by non-academics: e.g. licence needed
%\end{itemize}
\section{Availability of supporting data and materials}
%``The data set(s) supporting the results of this article is(are) available in the [repository name] repository, [cite unique persistent identifier].''
This manuscript, and its supporting data are available under a Creative Commons Attribution 4.0 International license.
Unprocessed FASTQ from the Illumina sequencing of the ten isolates are available at the European Nucleotide Archive, via the identifiers listed in Table \ref{tab:strains}, identifiers for the even and log community Illumina sequencing can be found in Table \ref{tab:illumina}.
%McIntyre \textit{et al}.\ PacBio assemblies were downloaded from \url{https://github.com/al-mcintyre/mCaller_analysis_scripts/tree/master/assemblies}. PacBio RSII and Sequel reads from McIntyre \textit{et al}.\ are available from the Sequence Read Archive via the identifiers listed in Table \ref{tab:strains}.
Both the raw signal, and basecalled FASTQ for our nanopore sequencing experiments are available at the European Nucleotide Archive, via the identifiers listed in Table \ref{tab:datasets}.
The \texttt{SPAdes}-assembled Illumina draft reference, and the collection of nanopore assemblies for each \texttt{wtdbg2} condition are linked to from our GitHub repository (\url{https://github.com/LomanLab/mockcommunity}), along with the \texttt{kraken2} database used for taxonomic classification of the assembled contigs.
%Our DNA extraction methods have been posted to protocols.io and can be accessed via \url{http://dx.doi.org/10.17504/protocols.io.x9tfr6n}.
Further updates (such as updated references, or new assemblies) will be made available through our project website \url{https://lomanlab.github.io/mockcommunity/}.
%Supplementary Python and R scripts used to conduct analyses are available in our Github repository: \url{https://github.com/LomanLab/mockcommunity}.
\section{Declarations}
%\subsection{List of abbreviations}
%If abbreviations are used in the text they should be defined in the text at first use, and a list of abbreviations should be provided in alphabetical order.
%\subsection{Ethical Approval (optional)}
%Not applicable
\subsection{Consent for publication}
Not applicable
\subsection{Competing Interests}
Cambridge Biosciences provided the ZymoBIOMICS Microbial Community Standard free of charge.
ST is an employee of Zymo Research Corporation.
NJ has received Oxford Nanopore Technologies (ONT) reagents free of charge to support his research programme.
NJ and JQ have received travel expenses to speak at ONT events.
NL has received an honorarium to speak at an ONT company meeting.
\subsection{Funding}
SN is funded by the Medical Research Foundation and the NIHR STOP-COLITIS project.
JQ is funded by the NIHR Surgical Reconstruction and Microbiology Research Centre.
The NIHR SRMRC is a partnership between The National Institute for Health Research, University Hospitals Birmingham NHS Foundation Trust, the University of Birmingham, and the Royal Centre for Defence Medicine.
NL is funded by an MRC Fellowship in Microbial Bioinformatics under the CLIMB project.
\subsection{Author's Contributions}
Conceptualization: NL, Methodology: NL JQ SN ST, Software: SN NL, Validation: SN NL, Formal analysis: SN NL, Investigation: NL JQ SN, Resources: NL ST, Data Curation: SN NL ST, Writing – original draft preparation: SN, Writing – review and editing: SN NL JQ ST, Visualization: SN NL, Supervision: NL, Project administration: NL, Funding acquisition: NL ST
%The individual contributions of authors to the manuscript should be specified in this section. Guidance and criteria for authorship can be found in our \href{https://academic.oup.com/gigascience/pages/editorial_policies_and_reporting_standards}{editorial policies}. We would recommend you follow some kind of standardised taxonomy like the \href{http://docs.casrai.org/CRediT}{CASRAI CRediT} (Contributor Roles Taxonomy).
\section{Acknowledgements}
We are grateful to Radoslaw Poplawski (University of Birmingham) for assistance with CLIMB virtual machines and file systems to support this research. We thank Divya Mirrington (Oxford Nanopore Technologies) for advice on PromethION library preparation and sequencing. We thank Hannah McDonnell at Cambridge Biosciences for providing the ZymoBIOMICS Microbial Community Standards. We thank Jared Simpson (Ontario Institute for Cancer Research), Matt Loose (University of Nottingham) and John Tyson (University of British Columbia) for useful discussions and advice. We thank Christopher Mason and Alexa McIntyre (Cornell University) for making PacBio data available ahead of publication.
%\section{Authors' information (optional)}
%You may choose to use this section to include any relevant information about the author(s) that may aid the reader's interpretation of the article, and understand the standpoint of the author(s). This may include details about the authors' qualifications, current positions they hold at institutions or societies, or any other relevant background information. Please refer to authors using their initials. Note this section should not be used to describe any competing interests.
%% Specify your .bib file name here, without the extension
\bibliography{paper-refs}
\end{document}