Statistics¶
PLS-DA¶
hide_deconv.statistic.run_plsda(data, sample_sheet, sample_id_col, cohort_col, out_path, datasets_to_map=[], labels_data_map=[])
¶
Runs a PLS-DA (PLS2) model and saves the corresponding plots
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
DataFrame
|
Estimated composition to be used for PLS-DA |
required |
sample_sheet
|
DataFrame
|
Samples sheet holding clinical metainformation |
required |
sample_id_col
|
str
|
Column name linking the sample sheet with the estimated compositions |
required |
cohort_col
|
str
|
Column name holding the different cohorts |
required |
out_path
|
Path
|
Path, where the created plots will be stored |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Estimated Scores |
Source code in src/hide_deconv/statistic/plsda.py
88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 | |
Cohort differences¶
hide_deconv.statistic.run_mann_whitney_u(bulks, sample_list, sample_id_col, cohort_col, celltypes_to_normalize_to=[])
¶
Performs a Man Whitney U Test with FDR correction for two cohorts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bulks
|
DataFrame
|
bulk Samples (either Celltype x Sample or Gene x Sample) |
required |
sample_list
|
DataFrame
|
sample list (samples x clinical variables) |
required |
sample_id_col
|
str
|
Name of the column, that links to the bulks |
required |
cohort_col
|
str
|
Name of the column containing the cohort identifiers. |
required |
celltypes_to_normalize_to
|
list = []
|
Cell types to exclude before sample wise renormalization. |
[]
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame containing the Mean, Standard Deviation, p-Value and adjusted p-Value |
Source code in src/hide_deconv/statistic/mann_whitney_u.py
22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 | |
hide_deconv.statistic.run_kruskal_wallis(bulks, sample_list, sample_id_col, cohort_col, celltypes_to_normalize_to=[])
¶
Performs a Kruskal Wallis Test with FDR correction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bulks
|
DataFrame
|
bulk Samples (either Celltype x Sample or Gene x Sample) |
required |
sample_list
|
DataFrame
|
sample list (samples x clinical variables) |
required |
sample_id_col
|
str
|
Name of the column, that links to the bulks |
required |
cohort_col
|
str
|
Name of the column containing the cohort identifiers. |
required |
celltypes_to_normalize_to
|
list = []
|
Cell types to exclude before sample wise renormalization. |
[]
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame containing p-Value and adjusted p-Value |
Source code in src/hide_deconv/statistic/kruskal_wallis.py
20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 | |
hide_deconv.statistic.run_dunn(kruskal_results, bulks, sample_list, sample_id_col, cohort_col, sign_level=0.05)
¶
Performs a Posthoc Dunn Test on significant Kruskal Wallis results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
kruskal_results
|
DataFrame
|
Results of obtained by the run_kruskal_wallis() method |
required |
bulks
|
DataFrame
|
bulk Samples (either Celltype x Sample or Gene x Sample) |
required |
sample_list
|
DataFrame
|
sample list (samples x clinical variables) |
required |
sample_id_col
|
str
|
Name of the column, that links to the bulks |
required |
cohort_col
|
str
|
Name of the column containing the cohort identifiers. |
required |
sign_level
|
float = 0.05
|
Float, below which results are considered as significant. |
0.05
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame containing the results of the Dunn Test. |
Source code in src/hide_deconv/statistic/posthoc_dunn.py
22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 | |
Clustering¶
hide_deconv.statistic.run_clustering(data, is_bulk=False)
¶
Performs a clustering using greedy modular communities of the entered bulk or composition data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
DataFrame
|
Dataframe containing either bulks or composition (celltype/genes x samples) |
required |
is_bulk
|
bool = False
|
If set to true, uses correlation as distance measure for the neighborhood calculation. |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame containing the sample ids and the assigned clusters |
Source code in src/hide_deconv/statistic/clustering.py
15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 | |
Survival analysis¶
hide_deconv.statistic.run_cox_regression(bulks, sample_sheet, sample_id_col, time_col, event_col, covariates)
¶
Perform Cox Regression for each cell type composition.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bulks
|
DataFrame
|
Cell type compositions (either Celltype x Sample or Gene x Sample) |
required |
sample_sheet
|
DataFrame
|
sample list (samples x clinical variables) |
required |
sample_id_col
|
str
|
Name of the column, that links to the bulks |
required |
time_col
|
str
|
Name of the column with survival times |
required |
event_col
|
str
|
Name of column that indicates event (0: censored, 1: event) |
required |
covariates
|
list[str]
|
List of covariate column names |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with estimated impact on survival for each cell type |
Source code in src/hide_deconv/statistic/survival_analysis.py
31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 | |
Differential expression¶
hide_deconv.statistic.pydeseq2_preprocess(bulk, sample_sheet, sample_id_col, condition_col, covariates)
¶
Prepare bulk and sample metainfo for PyDESeq2.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bulk
|
DataFrame
|
Bulk RNA-seq file (genes x samples) |
required |
sample_sheet
|
DataFrame
|
Sample sheet containing the metainformation on samples (samples x info) |
required |
sample_id_col
|
str
|
Column name of the sample sheet containing the sample ids of the bulk file |
required |
condition_col
|
str
|
Column name of the sample sheet containing the conditions |
required |
covariates
|
list[str] | None
|
List of column names containing possible covariates |
required |
Returns:
| Type | Description |
|---|---|
tuple[DataFrame, DataFrame]
|
Counts, Metainformation to be used with the run_pydeseq2 function |
Source code in src/hide_deconv/statistic/pydeseq2.py
19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 | |
hide_deconv.statistic.run_pydeseq2(bulk, metadata, condition_col, tested_condition, reference_condition, covariates, out_path)
¶
Run PyDESeq2
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bulk
|
DataFrame
|
Raw count bulk RNA-seq data (Samples x Genes) |
required |
metadata
|
DataFrame
|
Sample metainformation |
required |
condition_col
|
str
|
Name of the column in metadata holding the condition |
required |
tested_condition
|
str
|
Name of the condition that will be tested |
required |
reference_condition
|
str
|
Name of the condition that is used as reference |
required |
covariates
|
list[str]
|
Column names of covariates to include |
required |
out_path
|
Folder path, where all results and plots are stored
|
|
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pydeseq2 result dataframe |
Source code in src/hide_deconv/statistic/pydeseq2.py
87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 | |