Bio-ToolBox
view release on metacpan or search on metacpan
scripts/correlate_position_data.pl view on Meta::CPAN
=item --norm [rank|sum]
Optionally define a method of normalizing the scores between the
reference and test data sets prior to calculating the correlation.
Two methods are currently supported: "rank" converts all values
to rank values (the mean rank is reported for identical values)
and essentially calculating a Spearman's rank correlation, while
"sum" scales all values so that the absolute sums are identical.
Normalization occurs after missing or zero values are interpolated.
The default is no normalization.
=item --force_strand
If enabled, a strand orientation will be enforced when determining the
optimal shift. This does not affect the correlation calculation, only
the direction of the reported shift. This requires the presence of a
data column in the input file with strand information. The default is
no enforcement of strand.
=back
=head2 General options
=over 4
=item --gz
Specify whether (or not) the output file should be compressed with gzip.
=item --version
Print the version number.
=item --help
Display this POD documentation.
=back
=head1 DESCRIPTION
This program will calculate statistics between the positioned scores of
two different datasets over a window from an annotated feature or
chromosomal segment. These statistics will help determine whether the
positions or distribution of scores across the window vary or underwent
a positional shift between a test and a reference dataset. For example,
if the enrichment of nucleosome signal from a ChIP experiment shifts in
genomic position, indicating a change in nucleosome position.
Two statistics may be calculated. First, it will calculate a a Pearson
linear correlation coefficient (r value) between the datasets (default).
Additionally, an ANOVA analysis may be performed between the datasets and
generate a P-value.
By default, the correlation is determined between the data points
collected over the entire length of the feature. Alternatively, a
radius and reference point (default is midpoint) may be provided
that sets the window for collecting scores and calculating a correlation.
In general, to ensure a more reliable Pearson value, fragment ChIP or
nucleosome coverage should be used rather than point (start or midpoint)
data, as it will give more reliable results. Fragment coverage is more
akin to smoothened data and gives better results than interpolated point
data.
Normalized read-depth data should be used when possible. If necessary,
Values can be normalized using one of two methods. The values may be
converted to rank positions (compare to Kendall's tau), or scaled such
that the absolute sum values are equal (for example, when working with
sequence tag read counts).
In addition to calculating a correlation coefficient, an optimal shift
may also be calculated. This essentially shifts the data, 1 bp at a time,
in order to identify a shift that would produce a higher correlation. In
other words, what amount of movement to the left or right would make the
test data look like the reference data? The window is shifted from -2
radius to +2 radius relative to the reference point, and the highest
correlation is reported along with the shift value that generated it.
=head1 AUTHOR
Timothy J. Parnell, PhD
Dept of Oncological Sciences
Huntsman Cancer Institute
University of Utah
Salt Lake City, UT, 84112
This package is free software; you can redistribute it and/or modify
it under the terms of the Artistic License 2.0.
( run in 1.466 second using v1.01-cache-2.11-cpan-364913b4093 )