COSG is a cosine similarity-based method for more accurate and scalable marker gene identification.
- COSG is a general method for cell marker gene identification across different data modalities, e.g., scRNA-seq, scATAC-seq, and spatially resolved transcriptome data.
- Marker genes or genomic regions identified by COSG are more indicative and with greater cell-type specificity.
- COSG is ultrafast for large-scale datasets and is capable of identifying marker genes for one million cells in less than two minutes.
The method and benchmarking results are described in Dai et al. (2022).
Additionally, the R version of COSG is available here.
Note I: we released our Python toolkit, PIASO, in which some methods were built upon COSG.
Note II: we have also recently released PIASOmarkerDB for beta testing.
Note III: COSG is also available for online analysis via Galaxy platform.
Stable version (PyPI):
pip install cosgStable version (bioconda):
conda install -c conda-forge -c bioconda cosgDevelopment version:
pip install git+https://github.com/genecell/COSG.gitRelease v1.1.3 (August 24, 2026)
- Fixed:
run_cosg_cytome(layer='auto')normalized an already-normalized matrix a second time.autochoselog1pfor RNA from a table keyed on the modality alone, without looking at the data. Becausecytome.from_anndatastores whateveradata.Xheld under{modality}_counts— including a log-normalized matrix — any file written from normalized input hadlog1papplied twice, while the function is documented as equivalent tocosg.cosg(adata). Correlation with the in-memory reference stayed around 0.99 and specificity moved by up to 0.11: small enough to pass a glance, large enough to change which genes rank as markers.autonow probes the stored matrix and, when it is demonstrably not integer, uses the values as given and warns, naming the override. An explicitlayer=is obeyed untouched. - Fixed: the version had two sources of truth.
pyproject.tomlandcosg/__init__.pyeach carried a literal, and they agree only until a release bumps one of them.pyprojectnow declares the version dynamic and readscosg.__version__. - Interoperates with cytome 0.3.0. That release stops
from_anndatawriting a non-integeradata.Xto{modality}_counts, so newer files carry{modality}_data(or a name you chose) instead.layer='auto'recognises this: it reads the recorded_anndata_X_layer, uses that matrix, and says which one it picked. Where cytome recordsmatrix_meta.is_integer, that is preferred over probing the values. The minimum cytome version is unchanged at 0.2.3 — this release exists to read older files correctly, and requiring 0.3.0 would help none of them. - If you have run
run_cosg_cytome(layer='auto')on a cytome converted from a normalized AnnData, re-run it. Runs on raw-count files are unaffected, and any run that passed an explicitlayer=was always correct.
Release v1.1.2 (August 22, 2026)
- Fixed:
cosg.cosg(adata)raisedImportErroron a default install. The cytomeDatasetalias lookup was not guarded like the first one, so a missing optional cytome dependency produced an error instead of the documented empty-tuple fallback. Plain AnnData no longer requires cytome to be installed. - Fixed:
plt.cm.get_cmap, removed in matplotlib 3.9, replaced withplt.get_cmap.
Release v1.1.1 (August 15, 2026)
- Fixed the PyPI project page, which showed a single sentence instead of this README. The
readmefield pointed at an inline string rather than atREADME.rst, so the long description was never packaged. Code is identical to v1.1.0.
Release v1.1.0 (August 12, 2026)
- Added a streaming backend for
.cytomedatasets: marker detection reads the file in chunks, so peak memory does not scale with the number of cells. - The cytome streaming backend is an optional extra, on the same reasoning as scanpy above:
pip install 'cosg[cytome]'. Callingcosg.cosg()on a.cytomepath without it raises an error naming the extra rather than failing obscurely. - The default streaming path no longer needs PIASO.
layer='log1p'(the RNA default) is computed by COSG itself and gives identical numbers to the implementation it replaced. Onlylayer='infog'andlayer='tfidf'need PIASO, since those are PIASO normalizations; the error naming it now gives the correct install command,pip install piaso-tools. - Added a GPU path (CuPy), selected with
device='cpu' | 'gpu' | 'auto'on both the in-memory and streaming paths. cosg.cosg()is now a single polymorphic entry point. It dispatches on its first argument:- an
AnnDatatakes the in-memory path and writesadata.uns[key_added] - a
strorpathlib.Pathto a.cytomefile takes the streaming path and returns a dict - an open
CytomeDataset(fromcytome.open()) also takes the streaming path; the caller keeps ownership of it and it is not closed
- an
- Extended
batch_keyto the streaming, GPU and feature-batched paths. - Added
output_formatfor the streaming path:'dict','long'or'dense'. - Added
cosg.__version__, which previously raisedAttributeError. - scanpy is now an optional extra rather than a required dependency. Only
plotMarkerDotplotuses it, so a plainpip install cosgis ~60 MB and 9 packages lighter. Installpip install 'cosg[dotplot]'to keep that function; calling it without scanpy raises an error naming the extra. Cell-type ordering that previously usedscanpy.tl.dendrogramis now computed internally and reproduces it exactly. - The streaming default layer is resolved from the modality:
RNA/GAtolog1p,ATAC/tilestotfidf. - Behaviour change:
remove_lowly_expressednow defaults toTrueon the AnnData path, matching the streaming variant. Passremove_lowly_expressed=Falseto restore the previous behaviour. - Behaviour change: IQR normalisation is computed over all values per group, matching
iqrLogNormalize.
Release v1.0.4 (March 5, 2026)
- Added
plotMarkerStreamfor visualising marker gene specificity as a streamgraph. - Added
expressed_min_num_cells_in_target_group(default 3), which floors the expression threshold atmax(n_cells * expressed_pct, 3)so small clusters do not get an overly permissive cutoff. - Added input validation for
groupby,groups,groupscombined withbatch_key, andn_genes_user.
Release v1.0.3 (March 11, 2025)
- Fixed the incompatibility with multiple index columns of
adata.uns['cosg']['COSG']inadata.writefunction - Enhanced
plotMarkerDendrogramfunction with several new capabilities:- Implemented support for customized cell type-gene pairs
- Added color control for nodes and edges
- Added cell type filtering functionality
- Integrated support for curved edges in visualization
Release v1.0.2 (March 5, 2025)
- Added
plotMarkerDotplotandplotMarkerDendrogramfor enhanced marker gene visualization. - Introduced support for
batch_keyto compute cosine similarities separately across different batches. - Enabled calculation of normalized COSG scores for comparing gene expression specificity across cell types or datasets.
- Resolved a SciPy version deprecation issue related to
.Aattribute usage. - Fixed a DataFrame manipulation warning.
- Added verbosity control, allowing users to adjust log output levels.
Release v1.0.1 (June 15, 2021)
- First release in PyPI.
Run COSG:
import cosg
n_genes=30
groupby='CellTypes'
cosg.cosg(
adata,
key_added='cosg',
# use_raw=False, layer='log1p', ## e.g., if you want to use the log1p layer in adata
mu=100,
expressed_pct=0.1,
remove_lowly_expressed=True,
n_genes_user=n_genes,
groupby=groupby
)Draw the dot plot:
cosg.plotMarkerDotplot(
adata,
groupby=groupby,
top_n_genes=3,
key_cosg='cosg',
use_rep='X_pca', ## Change use_rep to the cell embeddings key you'd like to use
swap_axes=False,
standard_scale='var',
cmap='Spectral_r',
# save='test.pdf'
)Output the marker list as pandas dataframe:
marker_gene=pd.DataFrame(adata.uns['cosg']['names'])
marker_gene.head()You could also check the COSG scores:
marker_gene_scores=pd.DataFrame(adata.uns['cosg']['scores'])
marker_gene_scores.head()For questions about the code and tutorial, please contact Min Dai, dai@broadinstitute.org.
If COSG is useful for your research, please consider citing Dai et al. (2022).
COSG runs the same algorithm in several modes, selected by the available hardware:
| Mode | Speedup | Memory | Requirements |
|---|---|---|---|
| CPU chunked | 1.8x | ~19 GB | Default, any system |
| CPU legacy | 1.0x | ~38 GB | cpu_chunk_size=0 |
| GPU monolithic | 6.7x | 15 GB VRAM | 24+ GB GPU |
| GPU chunked | 4.7x | 0.9 GB | Any GPU (4+ GB) |
Benchmarked on Allen Institute Human Neocortex (148K cells x 30K genes).
All four modes read an in-memory AnnData, so peak memory still scales with
the size of the object. To keep memory bounded by the chunk size instead, read
straight from a file — see Cytome datasets.
COSG does not modify the input adata.X. It is safe to call
cosg.cosg() without copying adata first.
pip install cosg covers marker detection, plotMarkerDendrogram and
plotMarkerStream. Two features need extras:
pip install 'cosg[dotplot]' # plotMarkerDotplot (wraps scanpy.pl.dotplot) pip install 'cosg[gpu]' # the CuPy GPU path, device='gpu'
scanpy was a hard dependency through 1.0.4. It is now optional because only
plotMarkerDotplot uses it, and requiring it cost every install ~60 MB and
9 packages (statsmodels, seaborn, umap-learn, ...) for one plotting function.
Calling plotMarkerDotplot without it raises an error naming the extra.
Marker genes can be computed directly from a .cytome file, without loading
the matrix into memory and without going through AnnData:
pip install cytomeimport cosg
markers = cosg.cosg(
"atlas.cytome",
groupby="cell_type",
modality="RNA",
layer="counts",
n_genes_user=50,
)cosg.cosg() dispatches on its first argument. An AnnData takes the
in-memory path and writes adata.uns[key_added]; a path to a .cytome
file takes the streaming path, reads the file in chunks and returns a dict of
names, scores and groups_order. Peak memory is set by the chunk
size, not by the number of cells.
layer= names a matrix stored in the file. Omit it and the default is
resolved from the modality — RNA/GA to log1p, ATAC/tiles
to tfidf — which normalizes on the fly and therefore also needs
pip install piaso. Pass a stored layer, as above, to run with cosg and
cytome alone. output_format= selects dict, long or dense.
cytome is a single-file, SQLite-backed format that holds matrices, cell and feature metadata, embeddings and genomic fragments together: https://github.com/genecell/cytome