Preprocessing¶
Reference profiles¶
hide_deconv.preprocessing.create_reference(adata, celltype_col='cell_type')
¶
Create archetypal cell type references from a given AnnData object.
The function creates reference profiles from a given AnnData object by averaging over the gene expression profiles of each cell type. The used cell types are determined by the given cell type observation name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
adata
|
AnnData
|
Input expression data. |
required |
celltype_col
|
str
|
Column in adata.obs containing the cell type labels used for averaging over the gene expressions. |
"cell_type"
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
A gene x cell type pandas DataFrame containing the archetypal gene expression profiles of each cell type. |
Source code in src/hide_deconv/preprocessing/train_preprocessing.py
76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 | |
Simulated bulk data¶
hide_deconv.preprocessing.create_bulks(adata, n_bulks, n_cells_per_bulk, celltype_col='cell_type', seed=42, norm=False)
¶
Generate in silico bulk expression samples and their cell type compositions.
The function draws cells with replacement from the input AnnData object to create n_bulks bulk samples. For each bulk, expression counts are summed across the sampled cells and the corresponding cell type counts are tracked. Optionally, bulk expression is normalized to counts per million (CPM).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
adata
|
AnnData
|
Input AnnData object. |
required |
n_bulks
|
int
|
Number of bulks samples to simulate. |
required |
n_cells_per_bulk
|
int
|
Number of cells sampled per simulated bulk. |
required |
celltype_col
|
str
|
Column in adata.obs containing cell type labels. |
"cell_type"
|
seed
|
int
|
Random seed for reproducible simulations. |
42
|
norm
|
bool
|
If True each bulk is normalized to CPM. |
False
|
Returns:
| Type | Description |
|---|---|
tuple[DataFrame, DataFrame]
|
Y : gene x bulk expression profiles. C : celltype x bulk composition matrix |
Source code in src/hide_deconv/preprocessing/train_preprocessing.py
246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 | |
Domain transfer¶
hide_deconv.preprocessing.get_domain_transfer_factor(df1, df2)
¶
Calculate the domain transfer factor to align two bulks
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df1
|
DataFrame
|
First dataframe |
required |
df2
|
DataFrame
|
Second dataframe |
required |
Returns:
| Type | Description |
|---|---|
Series
|
Series containing the conversion factors for each gene |
Source code in src/hide_deconv/preprocessing/bulk_preprocessing.py
65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 | |
Gene and bulk utilities¶
hide_deconv.preprocessing.get_common_genes(adata, bulk, remove_zero_median=True)
¶
Return genes present in both single-cell and bulk expression data.
The function computes the intersection between gene names in adata.var_names and the bulk expressions.
Additionally remove genes, which have a median of 0 in the bulk, as these can influence domain transfer.
Gene Symbols that are duplicate will be added together (Preferentably to use EnsemblIDs).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
adata
|
AnnData
|
AnnData object containing single-cell expression data. Gene names are taken from adata.var_names |
required |
bulk
|
DataFrame
|
Bulk expression profiles with genes as index and samples as columns |
required |
remove_zero_median
|
bool = True
|
Remove genes that have a median expression of zero |
True
|
Returns:
| Type | Description |
|---|---|
list[str]
|
List of shared gene names |
Source code in src/hide_deconv/preprocessing/bulk_preprocessing.py
13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 | |
hide_deconv.preprocessing.combine_bulk_dataframes(data_frames)
¶
Combines a list of bulk RNA seq dataframes and corrects for batch effects.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data_frames
|
list[DataFrame]
|
List of bulk RNA seq dataframes that should be combined. |
required |
Returns:
| Type | Description |
|---|---|
tuple[DataFrame, DataFrame]
|
DataFrame, containing the bulk labels as columns and the merged genes as rows and DataFrame containing assignment to original batches. |
Source code in src/hide_deconv/preprocessing/bulk_preprocessing.py
94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 | |
AnnData utilities¶
hide_deconv.preprocessing.reduce_genes(adata, N, ct_col='cell_type')
¶
Reduce an AnnData object to the N most informative genes.
The function computes the mean expression per cell type and selects the genes with the highest variance between cell types. It returns a copy of the AnnData object restricted to the selected genes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
adata
|
AnnData
|
Input expression data. |
required |
N
|
int
|
Number of genes to select. |
required |
ct_col
|
str
|
Column in adata.obs containing the cell type labels used for variance calculation |
"cell_type"
|
Returns:
| Type | Description |
|---|---|
AnnData
|
A copy of the input AnnData object containing only the selected genes. |
Source code in src/hide_deconv/preprocessing/train_preprocessing.py
17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 | |
hide_deconv.preprocessing.create_hierarchy(adata, ct_col_sub, ct_col_higher)
¶
Create hiearchy mapping matrices between subtypes and higher-level cell types.
The function constructs for each column in ct_col_higher a binary projection matrix that maps each subtype in ct_col_sub to its corresponding higher-level cell type. Each returned matrix has higher-level cell types as rows and subtypes as columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
adata
|
AnnData
|
Input expression data containing cell annotations in adata.obs |
required |
ct_col_sub
|
str
|
Column in adata.obs containing the lower-level cell type labels. |
required |
ct_col_higher
|
list[str]
|
List of column names in adata.obs containing the higher-level cell type labels for which hiearchy matrices should be created. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, DataFrame]
|
Dictionary mapping each higher-level annotation column name to a pandas DataFrame with shape (n_higher_types, n_subtypes), where entries are 1 if a subtype belongs to a higher-level type and 0 otherwise. |
Source code in src/hide_deconv/preprocessing/train_preprocessing.py
126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 | |
hide_deconv.preprocessing.train_test_split_adata(adata, celltype_col='cell_type', train_frac=0.5, seed=42)
¶
Split an AnnData object into a train and test subset.
The function splits the cells of each cell type independently into a training and a test subset. Samples are shuffled before splitting. For cell types with at least two samples the split ensures that both train and test receive at least one sample.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
adata
|
AnnData
|
Input AnnData object to split. |
required |
celltype_col
|
str
|
Column in adata.obs containing the cell type labels used for splitting |
"cell_type"
|
train_frac
|
float
|
Fraction of samplesper cell type assigned to the training split. |
0.5
|
seed
|
int
|
Random seed used for shuffling before the split. |
42
|
Returns:
| Type | Description |
|---|---|
tuple[AnnData, AnnData]
|
A tuple containing the training AnnData object and the test AnnData object. |
Source code in src/hide_deconv/preprocessing/train_preprocessing.py
184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 | |
hide_deconv.preprocessing.get_adata_info(ad_file)
¶
Returns a summary of a AnnData file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ad_file
|
str
|
Path to AnnData file. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, object]
|
Dictionary containing multiple metrics. |
Source code in src/hide_deconv/preprocessing/train_preprocessing.py
330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 | |