Skip to content

Preprocessing

Reference profiles

hide_deconv.preprocessing.create_reference(adata, celltype_col='cell_type')

Create archetypal cell type references from a given AnnData object.

The function creates reference profiles from a given AnnData object by averaging over the gene expression profiles of each cell type. The used cell types are determined by the given cell type observation name.

Parameters:

Name Type Description Default
adata AnnData

Input expression data.

required
celltype_col str

Column in adata.obs containing the cell type labels used for averaging over the gene expressions.

"cell_type"

Returns:

Type Description
DataFrame

A gene x cell type pandas DataFrame containing the archetypal gene expression profiles of each cell type.

Source code in src/hide_deconv/preprocessing/train_preprocessing.py
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
def create_reference(
    adata: ad.AnnData, celltype_col: str = "cell_type"
) -> pd.DataFrame:
    """
    Create archetypal cell type references from a given AnnData object.

    The function creates reference profiles from a given AnnData object by
    averaging over the gene expression profiles of each cell type.
    The used cell types are determined by the given cell type observation name.

    Parameters
    ----------
    adata : anndata.AnnData
        Input expression data.
    celltype_col : str, default="cell_type"
        Column in adata.obs containing the cell type labels used for averaging over the gene expressions.

    Returns
    -------
    pd.DataFrame
        A gene x cell type pandas DataFrame containing the archetypal gene expression profiles of each cell type.

    """

    celltypes = adata.obs[celltype_col].astype(str).values

    X = adata.X
    if sp.issparse(X):
        X = X.tocsr()
    else:
        X = np.asarray(X)

    unique_ct = np.array(sorted(pd.unique(celltypes)))
    ref = np.zeros((adata.n_vars, len(unique_ct)), dtype=np.float64)

    for j, ct in enumerate(unique_ct):
        mask = celltypes == ct
        if mask.sum() == 0:
            continue
        if sp.issparse(X):
            ref[:, j] = np.asarray(X[mask].mean(axis=0)).ravel()
        else:
            ref[:, j] = X[mask].mean(axis=0)

    return pd.DataFrame(ref, index=adata.var_names, columns=unique_ct)

Simulated bulk data

hide_deconv.preprocessing.create_bulks(adata, n_bulks, n_cells_per_bulk, celltype_col='cell_type', seed=42, norm=False)

Generate in silico bulk expression samples and their cell type compositions.

The function draws cells with replacement from the input AnnData object to create n_bulks bulk samples. For each bulk, expression counts are summed across the sampled cells and the corresponding cell type counts are tracked. Optionally, bulk expression is normalized to counts per million (CPM).

Parameters:

Name Type Description Default
adata AnnData

Input AnnData object.

required
n_bulks int

Number of bulks samples to simulate.

required
n_cells_per_bulk int

Number of cells sampled per simulated bulk.

required
celltype_col str

Column in adata.obs containing cell type labels.

"cell_type"
seed int

Random seed for reproducible simulations.

42
norm bool

If True each bulk is normalized to CPM.

False

Returns:

Type Description
tuple[DataFrame, DataFrame]

Y : gene x bulk expression profiles. C : celltype x bulk composition matrix

Source code in src/hide_deconv/preprocessing/train_preprocessing.py
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
def create_bulks(
    adata: ad.AnnData,
    n_bulks: int,
    n_cells_per_bulk: int,
    celltype_col: str = "cell_type",
    seed: int = 42,
    norm: bool = False,
) -> tuple[pd.DataFrame, pd.DataFrame]:
    """
    Generate in silico bulk expression samples and their cell type compositions.

    The function draws cells with replacement from the input AnnData object to create
    n_bulks bulk samples. For each bulk, expression counts are summed across the
    sampled cells and the corresponding cell type counts are tracked. Optionally, bulk
    expression is normalized to counts per million (CPM).

    Parameters
    ----------
    adata : ad.AnnData
        Input AnnData object.
    n_bulks : int
        Number of bulks samples to simulate.
    n_cells_per_bulk : int
        Number of cells sampled per simulated bulk.
    celltype_col : str, default="cell_type"
        Column in adata.obs containing cell type labels.
    seed : int, default=42
        Random seed for reproducible simulations.
    norm : bool, default=False
        If True each bulk is normalized to CPM.

    Returns
    -------
    tuple[pd.DataFrame, pd.DataFrame]
        Y : gene x bulk expression profiles.
        C : celltype x bulk composition matrix
    """

    rng = np.random.default_rng(seed)

    celltypes = adata.obs[celltype_col].astype(str).values
    if hasattr(adata.X, "toarray"):
        expr = adata.X.toarray()
    elif hasattr(adata.X, "todense"):
        expr = adata.X.todense()
    else:
        expr = np.array(adata.X[:])

    n_cells, n_genes = expr.shape

    unique_celltypes = sorted(pd.unique(celltypes))
    n_celltypes = len(unique_celltypes)
    celltype_to_idx = {ct: i for i, ct in enumerate(unique_celltypes)}

    Y = np.zeros((n_genes, n_bulks), dtype=np.float32)
    C = np.zeros((n_celltypes, n_bulks), dtype=np.float32)

    for b in range(n_bulks):
        idx = rng.choice(n_cells, size=n_cells_per_bulk, replace=True)
        sampled_expr = expr[idx]
        sampled_celltypes = celltypes[idx]

        Y[:, b] = sampled_expr.sum(axis=0)

        for ct in sampled_celltypes:
            C[celltype_to_idx[ct], b] += 1

    if norm:
        Y = (Y * 1e6) / Y.sum(axis=0)  # cpm

    Y = pd.DataFrame(
        Y, index=adata.var_names, columns=[f"bulk_{i}" for i in range(n_bulks)]
    )
    C = pd.DataFrame(
        C, index=unique_celltypes, columns=[f"bulk_{i}" for i in range(n_bulks)]
    )
    C = C / C.sum(axis=0)

    return Y, C

Domain transfer

hide_deconv.preprocessing.get_domain_transfer_factor(df1, df2)

Calculate the domain transfer factor to align two bulks

Parameters:

Name Type Description Default
df1 DataFrame

First dataframe

required
df2 DataFrame

Second dataframe

required

Returns:

Type Description
Series

Series containing the conversion factors for each gene

Source code in src/hide_deconv/preprocessing/bulk_preprocessing.py
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
def get_domain_transfer_factor(df1: pd.DataFrame, df2: pd.DataFrame) -> pd.Series:
    """
    Calculate the domain transfer factor to align two bulks

    Parameters
    ----------
    df1 : pd.DataFrame
        First dataframe
    df2 : pd.DataFrame
        Second dataframe

    Returns
    -------
    pd.Series
        Series containing the conversion factors for each gene
    """
    # \sum_g df1_g = \alpha \sum_g df2_g

    mean_df1 = df1.median(axis=1)
    mean_df2 = df2.median(axis=1)

    alpha = mean_df1 / mean_df2

    return alpha

Gene and bulk utilities

hide_deconv.preprocessing.get_common_genes(adata, bulk, remove_zero_median=True)

Return genes present in both single-cell and bulk expression data.

The function computes the intersection between gene names in adata.var_names and the bulk expressions.

Additionally remove genes, which have a median of 0 in the bulk, as these can influence domain transfer.

Gene Symbols that are duplicate will be added together (Preferentably to use EnsemblIDs).

Parameters:

Name Type Description Default
adata AnnData

AnnData object containing single-cell expression data. Gene names are taken from adata.var_names

required
bulk DataFrame

Bulk expression profiles with genes as index and samples as columns

required
remove_zero_median bool = True

Remove genes that have a median expression of zero

True

Returns:

Type Description
list[str]

List of shared gene names

Source code in src/hide_deconv/preprocessing/bulk_preprocessing.py
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
def get_common_genes(
    adata: ad.AnnData, bulk: pd.DataFrame, remove_zero_median=True
) -> list[str]:
    """
    Return genes present in both single-cell and bulk expression data.

    The function computes the intersection between gene names in adata.var_names
    and the bulk expressions.

    Additionally remove genes, which have a median of 0 in the bulk, as these can influence domain transfer.

    Gene Symbols that are duplicate will be added together (Preferentably to use EnsemblIDs).

    Parameters
    ----------
    adata : ad.AnnData
        AnnData object containing single-cell expression data.
        Gene names are taken from adata.var_names

    bulk : pd.DataFrame
        Bulk expression profiles with genes as index and samples as columns

    remove_zero_median : bool = True
        Remove genes that have a median expression of zero

    Returns
    -------
    list[str]
        List of shared gene names

    """

    # Handle duplicate Gene Symbols, appear sometimes after converting from EnsemblIDs to GeneSymbol
    if bulk.index.has_duplicates:
        bulk = bulk.groupby(level=0).sum()

    if remove_zero_median:
        bulk_med = bulk.median(axis=1)
        zero_genes = bulk_med[bulk_med == 0].index
        if len(zero_genes) > 0:
            bulk = bulk.drop(index=zero_genes)

    genes_sc = set(adata.var_names)
    genes_bulk = set(bulk.index)
    common_genes = list(genes_sc.intersection(genes_bulk))

    return common_genes

hide_deconv.preprocessing.combine_bulk_dataframes(data_frames)

Combines a list of bulk RNA seq dataframes and corrects for batch effects.

Parameters:

Name Type Description Default
data_frames list[DataFrame]

List of bulk RNA seq dataframes that should be combined.

required

Returns:

Type Description
tuple[DataFrame, DataFrame]

DataFrame, containing the bulk labels as columns and the merged genes as rows and DataFrame containing assignment to original batches.

Source code in src/hide_deconv/preprocessing/bulk_preprocessing.py
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
def combine_bulk_dataframes(
    data_frames: list[pd.DataFrame],
) -> tuple[pd.DataFrame, pd.DataFrame]:
    """
    Combines a list of bulk RNA seq dataframes and corrects for batch effects.

    Parameters
    ----------
    data_frames : list[pd.DataFrame]
        List of bulk RNA seq dataframes that should be combined.
    Returns
    -------
    tuple[pd.DataFrame, pd.DataFrame]
        DataFrame, containing the bulk labels as columns and the merged genes as rows and DataFrame containing assignment to original batches.
    """

    for i in range(len(data_frames)):
        data_frames[i] = (
            data_frames[i] * 1e6 / data_frames[i].sum(axis=0)
        )  # CPM everything before reducing genes!

    common_genes = data_frames[0].index
    for i, df in enumerate(data_frames[1:]):
        common_genes = common_genes.intersection(df[df.median(axis=1) > 0].index)

    if len(common_genes) == 0:
        raise ValueError("No shared genes found between input bulk dataframes.")

    data_frames = [df.loc[common_genes].copy() for df in data_frames]

    batch = []
    for i, df in enumerate(data_frames):
        batch.extend([f"batch_{i}" for _ in range(df.shape[1])])

    if len(data_frames) > 1:
        ref = data_frames[0]
        ref_median = ref.median(axis=1)

        aligned = [ref]
        for df in data_frames[1:]:
            tgt_median = df.median(axis=1)
            alpha = (ref_median) / (tgt_median)
            aligned.append(df.mul(alpha, axis=0))

        data_frames = aligned

    combined = pd.concat([df for df in data_frames], axis=1)

    batch_df = pd.DataFrame(batch, columns=["batch"], index=combined.columns)

    return combined, batch_df

AnnData utilities

hide_deconv.preprocessing.reduce_genes(adata, N, ct_col='cell_type')

Reduce an AnnData object to the N most informative genes.

The function computes the mean expression per cell type and selects the genes with the highest variance between cell types. It returns a copy of the AnnData object restricted to the selected genes.

Parameters:

Name Type Description Default
adata AnnData

Input expression data.

required
N int

Number of genes to select.

required
ct_col str

Column in adata.obs containing the cell type labels used for variance calculation

"cell_type"

Returns:

Type Description
AnnData

A copy of the input AnnData object containing only the selected genes.

Source code in src/hide_deconv/preprocessing/train_preprocessing.py
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
def reduce_genes(adata: ad.AnnData, N: int, ct_col: str = "cell_type") -> ad.AnnData:
    """
    Reduce an AnnData object to the N most informative genes.

    The function computes the mean expression per cell type and selects
    the genes with the highest variance between cell types. It returns a
    copy of the AnnData object restricted to the selected genes.

    Parameters
    ----------
    adata : anndata.AnnData
        Input expression data.
    N : int
        Number of genes to select.
    ct_col : str, default="cell_type"
        Column in adata.obs containing the cell type labels used for variance calculation

    Returns
    -------
    anndata.AnnData
        A copy of the input AnnData object containing only the selected genes.

    """
    n_genes_total = adata.n_vars

    N = min(N, n_genes_total)

    ct = adata.obs[ct_col].astype(str).values
    celltypes = np.array(sorted(pd.unique(ct)))

    X = adata.X
    if sp.issparse(X):
        X = X.tocsr()
    else:
        X = np.asarray(X)

    mean_by_ct = np.zeros((len(celltypes), n_genes_total), dtype=np.float64)

    for i, c in enumerate(celltypes):
        rows = np.where(ct == c)[0]
        if rows.size == 0:
            continue
        if sp.issparse(X):
            mean_by_ct[i, :] = np.asarray(X[rows].mean(axis=0)).ravel()
        else:
            mean_by_ct[i, :] = X[rows, :].mean(axis=0)

    var_between = mean_by_ct.var(axis=0)

    top_idx = np.argsort(-var_between)[:N]
    top_genes = adata.var_names[top_idx]

    adata_red = adata[:, top_genes].copy()
    return adata_red

hide_deconv.preprocessing.create_hierarchy(adata, ct_col_sub, ct_col_higher)

Create hiearchy mapping matrices between subtypes and higher-level cell types.

The function constructs for each column in ct_col_higher a binary projection matrix that maps each subtype in ct_col_sub to its corresponding higher-level cell type. Each returned matrix has higher-level cell types as rows and subtypes as columns.

Parameters:

Name Type Description Default
adata AnnData

Input expression data containing cell annotations in adata.obs

required
ct_col_sub str

Column in adata.obs containing the lower-level cell type labels.

required
ct_col_higher list[str]

List of column names in adata.obs containing the higher-level cell type labels for which hiearchy matrices should be created.

required

Returns:

Type Description
dict[str, DataFrame]

Dictionary mapping each higher-level annotation column name to a pandas DataFrame with shape (n_higher_types, n_subtypes), where entries are 1 if a subtype belongs to a higher-level type and 0 otherwise.

Source code in src/hide_deconv/preprocessing/train_preprocessing.py
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
def create_hierarchy(
    adata: ad.AnnData, ct_col_sub: str, ct_col_higher: list[str]
) -> dict[str, pd.DataFrame]:
    """
    Create hiearchy mapping matrices between subtypes and higher-level cell types.

    The function constructs for each column in ct_col_higher a binary projection matrix
    that maps each subtype in ct_col_sub to its corresponding higher-level cell type.
    Each returned matrix has higher-level cell types as rows and subtypes as columns.

    Parameters
    ----------
    adata : anndata.AnnData
        Input expression data containing cell annotations in adata.obs
    ct_col_sub : str
        Column in adata.obs containing the lower-level cell type labels.
    ct_col_higher : list[str]
        List of column names in adata.obs containing the higher-level
        cell type labels for which hiearchy matrices should be created.

    Returns
    -------
    dict[str, pd.DataFrame]
        Dictionary mapping each higher-level annotation column name to a
        pandas DataFrame with shape (n_higher_types, n_subtypes), where
        entries are 1 if a subtype belongs to a higher-level type and
        0 otherwise.
    """

    obs = adata.obs.copy()

    sub = obs[ct_col_sub].astype(str)
    sub_types = sorted(sub.unique())

    A_dict = {}

    for col in ct_col_higher:
        high = obs[col].astype(str)
        map_df = pd.DataFrame({"sub": sub, "high": high}).drop_duplicates()

        sub_to_high = map_df.set_index("sub")["high"].to_dict()

        high_types = sorted(high.unique())
        A = pd.DataFrame(0, index=high_types, columns=sub_types, dtype=int)

        for s in sub_types:
            h = sub_to_high.get(s, None)
            if h is not None:
                A.loc[h, s] = 1

        A_dict[col] = A

    return A_dict

hide_deconv.preprocessing.train_test_split_adata(adata, celltype_col='cell_type', train_frac=0.5, seed=42)

Split an AnnData object into a train and test subset.

The function splits the cells of each cell type independently into a training and a test subset. Samples are shuffled before splitting. For cell types with at least two samples the split ensures that both train and test receive at least one sample.

Parameters:

Name Type Description Default
adata AnnData

Input AnnData object to split.

required
celltype_col str

Column in adata.obs containing the cell type labels used for splitting

"cell_type"
train_frac float

Fraction of samplesper cell type assigned to the training split.

0.5
seed int

Random seed used for shuffling before the split.

42

Returns:

Type Description
tuple[AnnData, AnnData]

A tuple containing the training AnnData object and the test AnnData object.

Source code in src/hide_deconv/preprocessing/train_preprocessing.py
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
def train_test_split_adata(
    adata: ad.AnnData,
    celltype_col: str = "cell_type",
    train_frac: float = 0.5,
    seed: int = 42,
) -> tuple[ad.AnnData, ad.AnnData]:
    """
    Split an AnnData object into a train and test subset.

    The function splits the cells of each cell type independently into a training
    and a test subset. Samples are shuffled before splitting. For cell types
    with at least two samples the split ensures that both train and test
    receive at least one sample.

    Parameters
    ----------
    adata : anndata.AnnData
        Input AnnData object to split.
    celltype_col : str, default="cell_type"
        Column in adata.obs containing the cell type labels used for splitting
    train_frac : float, default=0.5
        Fraction of samplesper cell type assigned to the training split.
    seed : int, default=42
        Random seed used for shuffling before the split.

    Returns
    -------
    tuple[ad.AnnData, ad.AnnData]
        A tuple containing the training AnnData object and the test AnnData object.

    """

    rng = np.random.default_rng(seed)
    ct = adata.obs[celltype_col].astype(str).values

    train_idx = []
    test_idx = []

    for celltype in np.unique(ct):
        idx = np.where(ct == celltype)[0]
        rng.shuffle(idx)

        n_train = int(np.floor(len(idx) * train_frac))

        if len(idx) >= 2:
            n_train = max(1, min(n_train, len(idx) - 1))

        train_idx.extend(idx[:n_train])
        test_idx.extend(idx[n_train:])

    train_idx = np.array(train_idx, dtype=int)
    test_idx = np.array(test_idx, dtype=int)

    adata_train = adata[train_idx].copy()
    adata_test = adata[test_idx].copy()

    return adata_train, adata_test

hide_deconv.preprocessing.get_adata_info(ad_file)

Returns a summary of a AnnData file.

Parameters:

Name Type Description Default
ad_file str

Path to AnnData file.

required

Returns:

Type Description
dict[str, object]

Dictionary containing multiple metrics.

Source code in src/hide_deconv/preprocessing/train_preprocessing.py
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
def get_adata_info(ad_file: str) -> dict[str, object]:
    """
    Returns a summary of a AnnData file.

    Parameters
    ----------
    ad_file : str
        Path to AnnData file.

    Returns
    -------
    dict[str, object]
        Dictionary containing multiple metrics.
    """

    adata = ad.read_h5ad(ad_file)

    obs = adata.obs.columns.tolist()
    var = adata.var.columns.tolist()
    n_cells, n_genes = adata.shape
    return {"obs": obs, "var": var, "n_cells": n_cells, "n_genes": n_genes}