CRISPRessoSea is a tool for processing genome editing from multiple pooled amplicon sequencing experiments to characterizing editing at on- and off-targets across experimental conditions. The tool accepts raw sequencing files as input, performs analysis at each site in each sample, performs statistical analysis of editing, and produces plots and reports detailing editing rates.
Installation | Running modes | Tutorial | Parameters
CRISPRessoSea can be installed using Bioconda.
conda config --add channels defaults
conda config --add channels bioconda
conda config --add channels conda-forge
conda create -y -n CRISPRessoSea bioconda::crispressosea bioconda::cas-offinder # Create CRISPRessoSea conda environment
conda activate CRISPRessoSea
Alternatively, CRISPRessoSea can be installed directly from github into an environment that contains CRISPResso2 and its dependencies. cas-offinder is optional but required for creating a guide info file from scratch.
conda config --add channels defaults
conda config --add channels bioconda
conda config --add channels conda-forge
conda create -y -n CRISPRessoSea bioconda::crispresso2 bioconda::cas-offinder # Create CRISPRessoSea conda environment
conda activate CRISPRessoSea
pip install git+https://github.com/clementlab/CRISPRessoSea.git
CRISPRessoSea operates in three primary running modes to streamline the preparation, processing, and replotting of large-scale pooled sequencing experiments to measure and compare CRISPR genome editing.
MakeGuideFile: This mode enables users to generate a comprehensive target file that includes both on- and off-target sites for one or more guide RNA sequences. Off-target sites are computationally predicted using Cas-OFFinder, which reports genomic coordinates, sequences, and mismatch counts for each potential target. This functionality is especially useful for designing pooled experiments to profile editing activity at predicted off-target sites, or for preparing analyses when the off-target sequences or locations are not already known. The output of this mode is a standardized target file compatible with CRISPRessoSea’s Process mode.
Process: This program will process a pool of pools given:
- a target information file with headers: Guide (Name of the on-target sequence guide), Sequence, PAM (PAM sequence), #MM (Number of mismatches), and Locus (formatted like chr1:+2345). Target names can be provided in a 'Target' column.
- a sample file with headers: Name, fastq_r1, fastq_r2 (optional for single-end reads). A column specifying the group can also be provided in this file.
- a reference genome. The path to the reference genome and bowtie2 indices is provided not including the trailing .fa
Replot: This program will replot data from a finished analysis performed using Process. This avoids reprocessing or rerunning any analyses but allows users to reorder or subset guides to be plotted. This program accepts as input:
- A modified file derived from an 'aggregated_stats_all.txt' output from a completed
Processrun.
Download the tutorial dataset, untar it, and change into that directory by running:
wget https://github.com/clementlab/CRISPRessoSea/raw/refs/heads/main/demo/make_demo/small_demo/small_demo.tar.gz
tar -xzf small_demo.tar.gz
cd CRISPRessoSea_demo
This tutorial includes a subset of data from Cicera et al. 2020 investigating the CTLA4_Site9 guide with sequence GGACTGAGGGCCATGGACACNGG. The complete dataset can be found at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA625995.
For ease of use, I've created a super-small genome that includes the genomic sequence of only the on-target and three off-by-2 off-targets. The genomic sequence is contained in 'demo_genome.fa' and bowtie2 indices for the genome were generated using bowtie2-build demo_genome.fa demo_genome.
You can create a guide info file (e.g. for designing a pooled experiment to profile off-targets) by running:
CRISPRessoSea MakeGuideFile --guide_seq GGACTGAGGGCCATGGACAC --pam NGG --guide_name CTLA4_site9 --max_mismatches 2 --genome_file demo_genome.fa
- The
--guide_seqparameter specify the on-target guide sequence and does not include the PAM sequence. - The
--guide_nameis for convenience, and all off-targets are annotated with this guide name (seeGuidecolumn in output below). - The
--pamand--max_mismatchesparameters are used for finding off-target locations. Here, for our small example we'll search for up to 2 mismatches, but in practice up to 4 or 5 mismatches are investigated. - The
--genome_fileparameter specifies the path to the genome file. Here we are using a super-small genome with only the on-target and three off-by-2 off-targets.
Note that running MakeGuideFile requires Cas-offinder for enumerating off-target sites.
This produces a guide info file CRISPRessoSea_MakeGuideFileOutput/CRISPRessoSea.guide_info.txt that contains the on- and off-target locations for the CTLA4_site9 guide.:
Guide Target Sequence PAM #MM Locus Mismatch_info
CTLA4_site9 CTLA4_site9_ON_CTLA4_site0_500 GGACTGAGGGCCATGGACAC GGG 0 CTLA4_site0:+500 Mismatches: 0
CTLA4_site9 CTLA4_site9_OB2_CTLA4_site1_500 GGACaGAGGGCCcTGGACAC AGG 2 CTLA4_site1:+500 Mismatches: 2
CTLA4_site9 CTLA4_site9_OB2_CTLA4_site2_500 GGAaTGAGGcCCATGGACAC TGG 2 CTLA4_site2:+500 Mismatches: 2
CTLA4_site9 CTLA4_site9_OB2_CTLA4_site3_500 GGACTGgGGGCCtTGGACAC AGG 2 CTLA4_site3:+500 Mismatches: 2
- The
Guidecolumn is the name of the on-target guide. This value is the same for all guides (because they all refer to the same on-target guide). Targets will be grouped byGuidevalue for plotting. - The
Targetcolumn specifies a customizable target name that will appear in reports and plots. Here,MakeGuideFileassigns each target a different name based on the Guide, the number of mismatches, and the chromosome and start position. In the super-small genome file, I named the chromosomes 'CTLA4_site0', 'CTLA4_site1', etc. - The
SequenceandPAMcolumns show the target sequence and PAM information. - The
#MMcolumn shows the number of mismatches between the on-target and the off-target. - The
Locuscolumn shows the genomic location of the target in the form Chromosome : Strand Position - The
Mismatch_infocolumn is not required, but includes text describing the number of mismatches and bulges (if any). Note that Cas-offinder 3 is required for enumerating off-targets with bulges.
If you ran MakeGuideFile, you can now use the identified off-targets to run in Process mode using the command:
CRISPRessoSea Process --sample_file samples.demo.txt --target_file CRISPRessoSea_MakeGuideFileOutput/CRISPRessoSea.guide_info.txt --genome_file demo_genome.fa
- The
--sample_fileparameter specifies a sample file specifying the names and fastq sequences of samples with columnsName,fastq_r1, and optionallyfastq_r2andgroup. - The
--target_fileparameter specifies a target file specifying the targets - in this case it is produced by the MakeGuideFile function. - The
--genome_fileparameter specifies the path to the genome file.
If you didn't run MakeGuideFile, you can use the guide_info file in the demo dataset (see samples.demo.txt and guides.demo.txt):
CRISPRessoSea Process --sample_file samples.demo.txt --target_file guides.demo.txt --genome_file demo_genome.fa
The Process function will produce an html report at CRISPRessoSea_output_on_samples.demo.txt/output_crispresso_sea.html, as well as an aggregated statistics file at CRISPRessoSea_output_on_samples.demo.txt/aggregated_stats_all.txt which can be modified and input for the Replot function below.
The aggregated_stats_all.txt file includes a row for each target. Columns specify information about each target (columns target_id, target_name, target_chr, target_pos, etc) as well as aggregated information from each sample (columns that end with _highest_a_g_pct, _highest_c_t_pct, _highest_indel_pct and _tot_reads).
If you'd like to change the order and name of guides, or add or change statistical tests, you can replot using the command:
CRISPRessoSea Replot --output_folder replot.output --reordered_stats_file replot_agg_stats.txt --reordered_sample_file replot_samples.txt --sig_method_parameters t_test,Control,Treated,0.05
- The
--output_folderparameter specifies where the replotted output will be produced. - The
--reordered_stats_fileparameter is a modified aggregated stats file in the format produced byProcess. - The
--reordered_sample_fileparameter is a modified sample file where samples can be reordered if necessary. - The
--sig_method_parametersparameter specifies the significance test to be applied in the form of: - none
- hard_cutoff,{cutoff}
- mean_diff,{group1},{group2},{cutoff}
- t_test,{group1},{group2},{alpha}
- mann_whitney,{group1},{group2},{alpha}
- neg_binomial,{group1},{group2},{alpha}
Here, we specify t_test,Control,Treated,0.05 which specifies the t_test comparing Control vs Treated with a significance threshold of 0.05
usage: CRISPRessoSea.py MakeGuideFile [-h] [-o OUTPUT_FOLDER] [-p FILE_PREFIX] -g GUIDE_SEQ [-gn GUIDE_NAME] [--pam PAM] -x GENOME_FILE [--max_mismatches MAX_MISMATCHES] [--max_dna_bulges MAX_DNA_BULGES] [--max_rna_bulges MAX_RNA_BULGES] [-v VERBOSITY] [--debug]
options:
-h, --help show this help message and exit
-o OUTPUT_FOLDER, --output_folder OUTPUT_FOLDER
Output folder (default: None)
-p FILE_PREFIX, --file_prefix FILE_PREFIX
File prefix for output files (default: CRISPRessoSea)
-g GUIDE_SEQ, --guide_seq GUIDE_SEQ
Guide sequence(s) to create a guide file for. Multiple guides may be separated by commas. (default: None)
-gn GUIDE_NAME, --guide_name GUIDE_NAME
Guide name(s) to use for the guide sequence(s). Multiple names may be separated by commas. (default: guide_0)
--pam PAM PAM sequence(s) to use for the guide sequence(s). (default: NGG)
-x GENOME_FILE, --genome_file GENOME_FILE
Bowtie2-indexed genome file - files ending in and .bt2 must be present in the same folder. (default: None)
--max_mismatches MAX_MISMATCHES
Maximum number of mismatches to allow in the discovered offtargets (default: 4)
--max_dna_bulges MAX_DNA_BULGES
Maximum number of DNA bulges to allow in the discovered offtargets. Note that Cas-OFFinder 3 is required to detect sites with bulges. (default: 0)
--max_rna_bulges MAX_RNA_BULGES
Maximum number of RNA bulges to allow in the discovered offtargets. Note that Cas-OFFinder 3 is required to detect sites with bulges. (default: 0)
-v VERBOSITY, --verbosity VERBOSITY
Verbosity level of output to the console (1-4) 4 is the most verbose (default: 3)
--debug Print debug information (default: False)
usage: CRISPRessoSea.py Process [-h] [-o OUTPUT_FOLDER] -t TARGET_FILE -s SAMPLE_FILE [--gene_annotations GENE_ANNOTATIONS] -x GENOME_FILE [-r REGION_FILE] [--aggregation_method AGGREGATION_METHOD] [-p N_PROCESSES]
[--crispresso_quantification_window_center CRISPRESSO_QUANTIFICATION_WINDOW_CENTER] [--crispresso_quantification_window_size CRISPRESSO_QUANTIFICATION_WINDOW_SIZE]
[--crispresso_base_editor_output] [--crispresso_default_min_aln_score CRISPRESSO_DEFAULT_MIN_ALN_SCORE] [--crispresso_plot_window_size CRISPRESSO_PLOT_WINDOW_SIZE]
[--crispresso_ignore_substitutions] [--alternate_alleles ALTERNATE_ALLELES] [--allow_unplaced_chrs] [--plot_only_complete_targets] [--min_amplicon_coverage MIN_AMPLICON_COVERAGE] [--sort_based_on_mismatch]
[--allow_target_match_to_other_region_loc] [--top_percent_cutoff TOP_PERCENT_CUTOFF] [--min_amplicon_len MIN_AMPLICON_LEN] [--fail_on_pooled_fail]
[--plot_group_order PLOT_GROUP_ORDER] [--sig_method_parameters SIG_METHOD_PARAMETERS] [-v VERBOSITY] [--debug] [--version]
options:
-h, --help show this help message and exit
-o OUTPUT_FOLDER, --output_folder OUTPUT_FOLDER
Output folder (default: None)
-t TARGET_FILE, --target_file TARGET_FILE
Target file - list of targets with one target per line and the headers ['Guide','Sequence','PAM','#MM','Locus'] (default: None)
-s SAMPLE_FILE, --sample_file SAMPLE_FILE
Sample file - list of samples with one sample per line with headers ['Name','fastq_r1','fastq_r2'] (default: None)
--gene_annotations GENE_ANNOTATIONS
Gene annotations file - a tab-separated .bed file with gene annotations with columns "chr" or "chrom", "start" or "txstart", and "end" or "txend" as well as "name" (default: None)
-x GENOME_FILE, --genome_file GENOME_FILE
Bowtie2-indexed genome file - files ending in and .bt2 must be present in the same folder. (default: None)
-r REGION_FILE, --region_file REGION_FILE
Region file - a tab-separated .bed file with regions to analyze with columns for 'chr', 'start', and 'end'. If not provided, regions will be inferred by read alignment. (default: None)
--aggregation_method AGGREGATION_METHOD
Method for aggregating analysis outputs. Options are 'mod_pct' (takes the CRISPResso modification percentage), 'max_a_g' (for A>G base editors), 'max_c_t' (for C>T base editors), 'max_indel' (for indels), or 'all' (all comparisons are reported). (default: mod_pct)
-p N_PROCESSES, --n_processes N_PROCESSES
Number of processes to use. Set to "max" to use all available processors. (default: 8)
--crispresso_quantification_window_center CRISPRESSO_QUANTIFICATION_WINDOW_CENTER
Center of quantification window to use within respect to the 3' end of the provided sgRNA sequence. Remember that the sgRNA sequence must be entered without the PAM. For cleaving nucleases, this is the predicted cleavage position. The default is -3 and is suitable for the Cas9 system. For alternate nucleases, other cleavage offsets may be appropriate, for example, if using Cpf1 this parameter would be set to 1. For base editors, this could be set to -17 to only include mutations near the 5' end of the sgRNA. (default: -3)
--crispresso_quantification_window_size CRISPRESSO_QUANTIFICATION_WINDOW_SIZE
Size (in bp) of the quantification window extending from the position specified by the '--cleavage_offset' or '--quantification_window_center' parameter in relation to the provided guide RNA sequence(s) (--sgRNA). Mutations within this number of bp from the quantification window center are used in classifying reads as modified or unmodified. A value of 0 disables this window and indels in the entire amplicon are considered. Default is 1, 1bp on each side of the cleavage position for a total length of 2bp. (default: 1)
--crispresso_base_editor_output
Outputs plots and tables to aid in analysis of base editor studies. (default: False)
--crispresso_default_min_aln_score CRISPRESSO_DEFAULT_MIN_ALN_SCORE
Default minimum homology score for a read to align to a reference amplicon. (default: 60)
--crispresso_plot_window_size CRISPRESSO_PLOT_WINDOW_SIZE
Defines the size of the window extending from the quantification window center to plot. Nucleotides within plot_window_size of the quantification_window_center for each guide are plotted. (default: 20)
--crispresso_ignore_substitutions
Ignores substitutions when calling reads as modified or unmodified, affecting the 'mod_pct' columns. By default, substitutions are considered when calling reads as modified or unmodified. (default: False)
--alternate_alleles ALTERNATE_ALLELES
Alternate alleles file to be passed to CRISPRessoPooled to override amplicon sequences for specific regions. (default: None)
--allow_unplaced_chrs
Allow regions on unplaced chromosomes (chrUn, random, etc). By default, regions on these chromosomes are excluded. If set, regions on these chromosomes will be included. (default: False)
--plot_only_complete_targets
Plot only targets with all values. If not set, all targets will be plotted. (default: False)
--min_amplicon_coverage MIN_AMPLICON_COVERAGE
Minimum number of reads to cover a location for it to be plotted. Otherwise, it will be set as NA (default: 10)
--sort_based_on_mismatch
Sort targets based on mismatch count. If true, the on-target will always be first (default: False)
--allow_target_match_to_other_region_loc
If true, targets can match to regions even if the target chr:start is not in that region (e.g. if the target sequence is found in that region). If false/unset, targets can only match to regions matching the target chr:start position. This flag should be set if the genome for guide design was not the same as the analysis genome. (default: False)
--top_percent_cutoff TOP_PERCENT_CUTOFF
The top percent of aligned regions (by region read depth) to consider in finding non-overlapping regions during demultiplexing. This is a float between 0 and 1. For example, if set to 0.2, the top 20% of regions (by read depth) will be considered. (default: 0.2)
--min_amplicon_len MIN_AMPLICON_LEN
The minimum length of an amplicon to consider in finding non-overlapping regions during demultiplexing. Amplicons shorter than this will be ignored. (default: 50)
--fail_on_pooled_fail
If true, fail if any pooled CRISPResso run fails. By default, processing will continue even if sub-CRISPResso commands fail. (default: False)
--plot_group_order PLOT_GROUP_ORDER
Order of the groups to plot (if None, the groups will be sorted alphabetically) (default: None)
--sig_method_parameters SIG_METHOD_PARAMETERS
Parameters for the significance method in the form of: none hard_cutoff,cutoff mean_diff,group1,group2,cutoff t_test,group1,group2,alpha mann_whitney,group1,group2,alpha neg_binomial,group1,group2,alpha (default: None)
-v VERBOSITY, --verbosity VERBOSITY
Verbosity level of output to the console (1-4) 4 is the most verbose (default: 3)
--debug Print debug information (default: False)
--version show program's version number and exit
usage: CRISPRessoSea.py Replot [-h] [-o OUTPUT_FOLDER] [-p FILE_PREFIX] -f REORDERED_STATS_FILE -s REORDERED_SAMPLE_FILE [--aggregation_method AGGREGATION_METHOD] [--fig_width FIG_WIDTH]
[--fig_height FIG_HEIGHT] [--title_fontsize TITLE_FONTSIZE] [--y_tick_fontsize Y_TICK_FONTSIZE] [--x_tick_fontsize X_TICK_FONTSIZE] [--nucleotide_fontsize NUCLEOTIDE_FONTSIZE]
[--legend_title_fontsize LEGEND_TITLE_FONTSIZE] [--seq_plot_ratio SEQ_PLOT_RATIO] [--plot_group_order PLOT_GROUP_ORDER] [--sig_method_parameters SIG_METHOD_PARAMETERS]
[--dot_plot_ylims DOT_PLOT_YLIMS] [--dot_plot_fig_width DOT_PLOT_FIG_WIDTH] [--dot_plot_fig_height DOT_PLOT_FIG_HEIGHT] [--heatmap_max_value HEATMAP_MAX_VALUE]
[--heatmap_min_value HEATMAP_MIN_VALUE] [-v VERBOSITY] [--debug] [--version]
options:
-h, --help show this help message and exit
-o OUTPUT_FOLDER, --output_folder OUTPUT_FOLDER
Output folder where output plots should be placed (default: None)
-p FILE_PREFIX, --file_prefix FILE_PREFIX
File prefix for output files (default: CRISPRessoSea)
-f REORDERED_STATS_FILE, --reordered_stats_file REORDERED_STATS_FILE
Reordered statistics file - made by reordering rows from aggregated_stats_all.txt (default: None)
-s REORDERED_SAMPLE_FILE, --reordered_sample_file REORDERED_SAMPLE_FILE
Reordered_sample_file - path to the sample file with headers: Name, group, fastq_r1, fastq_r2 (group is always optional, fastq_r2 is optional for single-end reads) (default: None)
--aggregation_method AGGREGATION_METHOD
Method for aggregating analysis outputs. Options are 'mod_pct' (takes the CRISPResso modification percentage), 'max_a_g' (for A>G base editors), 'max_c_t' (for C>T base editors), 'max_indel' (for indels), or 'all' (all comparisons are reported). (default: mod_pct)
--fig_width FIG_WIDTH
Width of the figure (default: 24)
--fig_height FIG_HEIGHT
Height of the figure (default: 24)
--title_fontsize TITLE_FONTSIZE
Font size for the title (default: 30)
--y_tick_fontsize Y_TICK_FONTSIZE
Font size for the y-axis tick labels (default: 16)
--x_tick_fontsize X_TICK_FONTSIZE
Font size for the x-axis tick labels (default: 16)
--nucleotide_fontsize NUCLEOTIDE_FONTSIZE
Font size for the nucleotide labels (default: 14)
--legend_title_fontsize LEGEND_TITLE_FONTSIZE
Font size for the legend and axis titles (default: 20)
--seq_plot_ratio SEQ_PLOT_RATIO
Ratio of the width of the sequence plot to the data plot (>1 means the seq plot is larger than the data plot) (default: 1)
--plot_group_order PLOT_GROUP_ORDER
Order of the groups to plot (if None, the groups will be sorted alphabetically) (default: None)
--sig_method_parameters SIG_METHOD_PARAMETERS
Parameters for the significance method in the form of: none hard_cutoff,cutoff mean_diff,group1,group2,cutoff t_test,group1,group2,alpha mann_whitney,group1,group2,alpha neg_binomial,group1,group2,alpha (default: None)
--dot_plot_ylims DOT_PLOT_YLIMS
Comma-separated min,max y-axis limits for the dot plot. If None, the y-axis limits will be set automatically. (default: None,None)
--dot_plot_fig_width DOT_PLOT_FIG_WIDTH
Width of the dot plot figures (default: 20)
--dot_plot_fig_height DOT_PLOT_FIG_HEIGHT
Height of the dot plot figures (default: 6)
--heatmap_max_value HEATMAP_MAX_VALUE
Maximum value for the heatmap color scale, where a value of 1 sets the max value color to 1% (if None, the maximum value will be determined automatically) (default: None)
--heatmap_min_value HEATMAP_MIN_VALUE
Minimum value for the heatmap color scale, where a value of 1 sets the min value color to 1% (if None, the minimum value will be determined automatically) (default: None)
-v VERBOSITY, --verbosity VERBOSITY
Verbosity level of output to the console (1-4) 4 is the most verbose (default: 3)
--debug Print debug information (default: False)
--version show program's version number and exit
CRISPRessoSea requires two main input files: a target file and a sample (experiment) file. Both can be provided as tab-delimited .txt or .xlsx files.
This file provides information about each target profiled in the amplicon sequencing experiment. Targets can include a combination of on- and off-targets. Targets will be grouped by the 'Guide' column in plots. Information for each target is provided as a row in this table.
If the off-target sequences or locations are not known, this file can be generated using the MakeGuideFile command above. Example: guides.demo.txt.
Required columns:
| Column | Description | Example |
|---|---|---|
| Guide | Name of the on-target guide (used to group on- and off-targets) | CTLA4_site9 |
| Sequence | Target DNA sequence (without PAM, either on-target or off-target) | GGACTGAGGGCCATGGACAC |
| PAM | Protospacer Adjacent Motif sequence of this target | GGG |
| #MM | Number of mismatches between the on-target and this target (0 for on-target) | 0 |
| Locus | Genomic location in the format chr:pos, chr:+pos, or chr:-pos for strand information. |
chr1:+203870828 |
Optional columns:
| Column | Description | Example |
|---|---|---|
| Target | Custom name for the target (used in plots/reports) | CTLA4_site9_ontarget |
| Mismatch_info | Human-readable description of mismatches and bulges (for reference only - not used in computation) | Mismatches: 0 |
Notes:
- Column names are case-insensitive.
- All required columns must be present.
- Additional columns will be ignored.
- Locus strand information is not required or used by the pipeline. The strand-containing format is accepted because it is produced by some upstream tools.
- The Locus information is used to assign targets to amplicon sequencing sites. This functionality can be disabled using the
--allow_target_match_to_other_region_locflag.
Example:
Guide Target Sequence PAM #MM Locus
EMX1 EMX_OT1 GGCCTTGTTCACATATAACT AGG 0 chr1:+2345
EMX1 EMX_OT2 GGCCTTGTTCACATATAACT AGG 1 chr1:+5678This file lists all sequencing samples to be analyzed in a CRISPRessoSea run. Each row specifies a sample name, the path to its sequencing data (FASTQ files), and optionally its experimental group and other metadata. Example: samples.demo.txt.
Required columns:
| Column | Description |
|---|---|
| Name | Unique name for the sample |
| fastq_r1 | Path to the R1 FASTQ file for this sample |
Optional columns:
| Column | Description |
|---|---|
| fastq_r2 | Path to the R2 FASTQ file (for paired-end reads) |
| group | Group label for the sample (used for grouping in plots/stats) |
Notes:
- The
Namecolumn must be unique for each sample. - The
fastq_r2column is optional and only needed for paired-end data. - The
groupcolumn is optional but recommended for group-based analysis and plotting. - Additional columns will be ignored.
Example:
Name fastq_r1 fastq_r2 group
Sample1 /path/sample1_R1.fq /path/sample1_R2.fq Control
Sample2 /path/sample2_R1.fq /path/sample2_R2.fq TreatedThis optional file allows you to override amplicon sequences for specific regions when using the --alternate_alleles parameter in Process mode. The file must be tab-delimited.
Format with header (Recommended): If a header is included, the file must contain the following columns:
region_nameortarget_name: The name of the target/region to override.reference_seqs: A comma-separated list of alternate allele sequences.
Format without header: If no header is provided, CRISPRessoSea assumes:
- The first column is the target/region name.
- The second column is a comma-separated list of alternate allele sequences.
Example:
target_name reference_seqs
CTLA4_site9_ontarget GGACTGAGGGCCATGGACAC,GGACTGAGGGCCATGGACACGGGSee also:
- The parsing logic for these files is implemented in
parse_target_info,parse_sample_file, andrewrite_alternate_allelesinCRISPRessoSea.py. - For more details, see the comments in the code or run with the
--helpflag.