Skip to content
Merged

Dev #16

Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions CRAN-SUBMISSION
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
Version: 0.0.0.9002
Date: 2025-04-07 17:20:29 UTC
SHA: a5d940843126c3de1efe60c349dde4c586afba28
Version: 0.0.3
Date: 2025-09-04 14:29:33 UTC
SHA: 66d7dcd409a24d5e1a6a5e6e04c7ae8b9257a8a2
9 changes: 5 additions & 4 deletions DESCRIPTION
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
Package: immundata
Title: A Unified Data Layer for Large-Scale Single-Cell, Spatial and Bulk Immunomics
Version: 0.0.4.9000
Version: 0.0.5
Authors@R:
person("Vadim I.", "Nazarov", , "support@immunomind.com", role = c("aut", "cre"),
comment = c(ORCID = "0000-0003-3659-2709"))
Expand All @@ -10,11 +10,11 @@ Description: Provides a unified data layer for single-cell, spatial and bulk
but for AIRR data, a.k.a. Adaptive Immune Receptor Repertoire, VDJ-seq, RepSeq, or
VDJ sequencing data.
License: Apache License (>= 2)
URL: https://immunomind.com/, https://github.com/immunomind/immundata, https://immunomind.github.io/immundata/
URL: https://immunomind.github.io/docs/, https://github.com/immunomind/immundata
BugReports: https://github.com/immunomind/immundata/issues
Encoding: UTF-8
Roxygen: list(markdown = TRUE)
RoxygenNote: 7.3.2
RoxygenNote: 7.3.3
Depends:
R (>= 4.1.0),
dplyr,
Expand All @@ -35,6 +35,7 @@ Imports:
utils
Suggests:
rmarkdown,
testthat (>= 3.0.0)
testthat (>= 3.0.0),
Seurat
Config/testthat/edition: 3
Config/testthat/parallel: true
1 change: 1 addition & 0 deletions NAMESPACE
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ export(annotate_barcodes)
export(annotate_chains)
export(annotate_immundata)
export(annotate_receptors)
export(annotate_seurat)
export(assert_receptor_schema)
export(filter_barcodes)
export(filter_immundata)
Expand Down
14 changes: 13 additions & 1 deletion R/core_immundata.R
Original file line number Diff line number Diff line change
Expand Up @@ -101,7 +101,7 @@ ImmunData <- R6Class(
private$.annotations
},

#' @field repertoires Get a vector of repertoire names after data aggregation with [agg_repertoires()]
#' @field repertoires Get a table of repertoires and their basic statistics.
repertoires = function() {
# TODO: cache repertoire table to memory if not very big?
if (!is.null(private$.repertoire_table)) {
Expand All @@ -111,6 +111,18 @@ ImmunData <- R6Class(
} else {
NULL
}
},

#' @field metadata Get a table of repertoires without their basic statistics.
metadata = function() {
if (!is.null(private$.repertoire_table)) {
private$.repertoire_table |>
select(c(imd_schema("repertoire"), self$schema_repertoire)) |>
collect() |>
arrange_at(vars(1))
} else {
NULL
}
}
)
)
9 changes: 6 additions & 3 deletions R/io_immundata_write.R
Original file line number Diff line number Diff line change
Expand Up @@ -34,8 +34,11 @@
#' along with the schema needed to reconstruct/interpret them.
#'
#' @return
#' Invisibly returns the input `idata` object. Its primary effect is creating
#' `metadata.json` and `annotations.parquet` files in the `output_folder`.
#' Invisibly returns the input `idata` object, saved to disk.
#' In other words, this allows you to create snapshots of the data in the
#' `output_folder`. Mind that by saving the object, you execute all the
#' stored computations, so this operations can take longer than expected.
#' Read more about snapshots on our website in the ["Concept" section](https://immunomind.github.io/docs/concepts/basics/immutability/).
#'
#' @seealso [read_immundata()] for loading the saved data, [read_repertoires()]
#' which uses this function internally, [ImmunData] class definition.
Expand Down Expand Up @@ -93,5 +96,5 @@ write_immundata <- function(idata, output_folder) {

cli::cli_alert_success("ImmunData files saved to [{output_folder}]")

invisible(idata)
invisible(read_immundata(output_folder))
}
66 changes: 66 additions & 0 deletions R/operations_external_annotate_seurat.R
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
#' @title Annotate a Seurat object from ImmunData (by barcode)
#'
#' @description
#' Copy selected columns from `idata$annotations` to Seurat metadata using
#' the cell barcode. This is the simplest way to transfer data from `immundata`
#' to Seurat object, e.g., for plotting data on UMAP.
#'
#' @param idata An [immundata::ImmunData] object.
#' @param sdata A Seurat object (cells are columns; barcodes are `colnames(sdata)`).
#' @param cols Character vector with column names to transfer from `idata$annotations`.
#' Typical choices: `"clonal_prop_bin"` or `"clonal_rank_bin"`.
#'
#' @return The updated Seurat object with new metadata columns.
#'
#' @seealso
#' [ImmunData], [SeuratObject::AddMetaData]
#'
#' @details
#' See functions `annotate_clonality_rank` and `annotate_clonality_prop` in `immunarch` package.
#'
#'
#' @examples
#' \dontrun{
#' # After annotating receptors:
#' idata <- annotate_clonality_prop(idata)
#'
#' # Transfer the clonality bin to Seurat and plot:
#' sdata <- annotate_seurat(idata, sdata, cols = "clonal_prop_bin")
#' Seurat::DimPlot(sdata, reduction = "umap", group.by = "clonal_prop_bin", shuffle = TRUE)
#'
#' # Alternative: rank bins
#' idata <- annotate_clonality_rank(idata, bins = c(10, 100))
#' sdata <- annotate_seurat(idata, sdata, cols = "clonal_rank_bin")
#' }
#'
#' @concept Annotation
#' @export
annotate_seurat <- function(idata,
sdata,
cols) {
checkmate::assert_r6(idata, "ImmunData")
checkmate::assert_class(sdata, classes = "Seurat")
checkmate::assert_character(cols, min.len = 1, any.missing = FALSE)

ann <- idata$annotations
bcsym <- immundata::imd_schema_sym("barcode")
bcname <- immundata::imd_schema("barcode")

missing_cols <- setdiff(c(bcname, cols), colnames(ann))
if (length(missing_cols) > 0) {
rlang::abort(
cli::format_inline("Column(s) {cli::col_cyan(missing_cols)} not found in idata$annotations.")
)
}

# TODO: mind that we use distinct() here - there could be edge cases
df <- ann |>
dplyr::select(barcode = !!bcsym, dplyr::all_of(cols)) |>
dplyr::distinct(.data$barcode, .keep_all = TRUE) |>
collect()

meta <- as.data.frame(df[, cols, drop = FALSE])
rownames(meta) <- df$barcode

Seurat::AddMetaData(sdata, metadata = meta)
}
2 changes: 2 additions & 0 deletions R/operations_utils.R
Original file line number Diff line number Diff line change
Expand Up @@ -129,6 +129,7 @@ annotate_tbl_distance <- function(tbl_data,
)

# TODO: Optimize it via SQL instead of cycles - if it is even needed...
# TODO: lump together multiple patterns in batches
for (i in seq_along(patterns)) {
p <- patterns[[i]]
col_name_out <- dist_cols[i]
Expand Down Expand Up @@ -188,6 +189,7 @@ annotate_tbl_distance <- function(tbl_data,
# 4) precompute sequence length before (!) any filtering, on data loading, and don't compute it here

# TODO: max dist. Left join - compute. Right join - filter

if (is.na(max_dist)) {
uniq <- uniq |>
as_duckdb_tibble()
Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,9 +28,9 @@
</div>

<p align="center">
<a href="https://immunomind.github.io/docs/tutorials/single-cell/">Tutorials</a>
<a href="https://immunomind.github.io/docs/tutorials/single_cell/">Tutorials</a>
|
<a href="https://immunomind.github.io/immundata/reference">API reference</a>
<a href="https://immunomind.github.io/immundata/reference/">API reference</a>
|
<a href=https://immunomind.github.io/docs/>Ecosystem</a>
|
Expand Down Expand Up @@ -1056,7 +1056,7 @@ ggplot2::ggplot(data = clonal_space_homeo) + geom_col(aes(x = Tissue, y = occupi
## 🧩 Use Cases

> [!TIP]
> Tutorial on `immundata` + `immunarch` is available [on the ecosystem website](https://immunomind.github.io/docs/tutorials/single-cell/).
> Tutorial on `immundata` + `immunarch` is available [on the ecosystem website](https://immunomind.github.io/docs/tutorials/single_cell/).
>
> Read the previous section about the analysis.
>
Expand Down
10 changes: 0 additions & 10 deletions cran-comments.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,2 @@
## R CMD check results

0 errors | 0 warnings | 1 note

* This is another try to submit `immundata`.
I fixed all the notes and warnings related to the package. There are some notes left
pointing to the potential misspellings of the terms in the DESCRIPTION, but those are
correct.

The CRAN check also points out that there is no `read_parquet_duckdb`, but it passes tests on my machine (I know, I know...),
and I added the reference to the package, and the resultant documentation links work. I hope this is fine.
Thank you!
4 changes: 3 additions & 1 deletion man/ImmunData.Rd

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

46 changes: 46 additions & 0 deletions man/annotate_seurat.Rd

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

5 changes: 2 additions & 3 deletions man/immundata-package.Rd

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

7 changes: 5 additions & 2 deletions man/write_immundata.Rd

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading