Visualize Sequence Embeddings: Apply t-SNE and UMAP in R

Visualize Sequence Embeddings: Apply t-SNE and UMAP in R

Biological research generates an explosion of high-dimensional data, particularly with the advent of advanced sequencing and 'omics' technologies. Decoding meaning from these complex datasets, such as protein or DNA sequence embeddings derived from deep learning models, poses a significant analytical challenge. Traditional visualization methods falter when confronted with hundreds or thousands of dimensions, obscuring the intrinsic patterns and relationships crucial for biological discovery. This article activates strategies to


Analyze biological datasets with R for statistics, visualization, and inference


by deploying two potent dimensionality reduction techniques: t-Distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP). We engineer a path to transform abstract numerical vectors into interpretable visual landscapes, revealing hidden clusters, subtle gradients, and functional insights within your biological sequences. Prepare to forge a powerful visual narrative from your high-dimensional biological data in R, gaining unparalleled leverage in your research.

Deconstructing High-Dimensional Biological Data: Forge Clarity from Complexity

Deconstructing High-Dimensional Biological Data: Forge Clarity from Complexity

High-dimensional biological data represents a frontier for discovery, yet its inherent complexity often obscures the underlying biological truths. Sequence embeddings, whether derived from deep learning models for proteins like ProtTrans or DNA sequences, condense intricate information into vectors with hundreds or even thousands of dimensions. Navigating this vast landscape requires specialized tools. Without effective dimensionality reduction, discerning patterns, identifying distinct biological states, or visualizing relationships becomes an insurmountable task. We activate t-SNE and UMAP as our primary instruments to overcome this hurdle, transforming abstract numerical representations into intuitive 2D or 3D visual maps. These techniques empower us to forge clarity from complexity, revealing the structural organization and functional divergences within our biological datasets.

Traditional methods like Principal Component Analysis (PCA) excel at capturing global variance, but often struggle to preserve the intricate local relationships crucial for distinguishing fine-grained biological clusters. Sequence embeddings frequently exhibit highly non-linear structures, which PCA may linearize, thereby distorting the true proximity of biologically similar entities. This necessitates algorithms capable of unraveling these complex manifolds. Our strategic approach centers on leveraging t-SNE and UMAP for their robust capabilities in non-linear dimensionality reduction. They are engineered to project high-dimensional data into a lower-dimensional space while maximally preserving the neighborhood structure, ensuring that sequences that are biologically similar in the high-dimensional space remain proximate in the visual representation. This foundational step is critical for any subsequent statistical analysis or hypothesis generation. We begin by setting up our R environment, ensuring all necessary computational modules are activated, and generating a representative synthetic dataset to demonstrate these powerful transformation processes.

# Install and load necessary R packages
# We validate package presence and install only if missing to optimize workflow.
if (!requireNamespace("Rtsne", quietly = TRUE)) install.packages("Rtsne")
if (!requireNamespace("umap", quietly = TRUE)) install.packages("umap")
if (!requireNamespace("ggplot2", quietly = TRUE)) install.packages("ggplot2")
if (!requireNamespace("dplyr", quietly = TRUE)) install.packages("dplyr")

library(Rtsne)
library(umap)
library(ggplot2)
library(dplyr)

# --- Simulate High-Dimensional Sequence Embeddings --- 
# In a real-world biological pipeline, 'sequence_embeddings' would originate 
# from models like ProtTrans, ESM, or k-mer frequency vectors. 
# Here, we engineer a synthetic dataset with two distinct clusters to exemplify 
# the capabilities of t-SNE and UMAP in separation and visualization.
set.seed(123) # Activate reproducibility
num_sequences <- 500
embedding_dimension <- 128 # A common dimension for biological sequence embeddings

# Forge Cluster 1: Represents one biological group (e.g., protein family A)
cluster1_data <- matrix(rnorm(num_sequences/2 * embedding_dimension, mean = 0, sd = 1),
                        nrow = num_sequences/2, ncol = embedding_dimension)
cluster1_labels <- rep("Enzyme_Family_A", num_sequences/2)

# Forge Cluster 2: Represents a distinct biological group (e.g., protein family B)
cluster2_data <- matrix(rnorm(num_sequences/2 * embedding_dimension, mean = 7, sd = 1),
                        nrow = num_sequences/2, ncol = embedding_dimension)
cluster2_labels <- rep("Enzyme_Family_B", num_sequences/2)

# Combine these engineered clusters into a single embedding matrix
sequence_embeddings_matrix <- rbind(cluster1_data, cluster2_data)
sequence_labels <- c(cluster1_labels, cluster2_labels)

# Store embeddings and labels in a data frame for streamlined analysis and visualization
embeddings_df <- as.data.frame(sequence_embeddings_matrix)
embeddings_df$label <- factor(sequence_labels) # Convert labels to factor for categorical plotting

# Validate the structure of our simulated data
print("Dimensions of the synthetic embedding matrix:")
print(dim(sequence_embeddings_matrix))
print("First 5 rows of the embedding data frame with labels:")
print(head(embeddings_df))
Activating t-SNE for Local Structure Discovery: Parameters and Precision

Activating t-SNE for Local Structure Discovery: Parameters and Precision

t-SNE stands as a formidable instrument for unmasking the local structure within high-dimensional biological sequence embeddings. Its core principle revolves around converting high-dimensional Euclidean distances between data points into conditional probabilities that represent similarities. It then minimizes the Kullback-Leibler divergence between these probability distributions in both the high-dimensional and low-dimensional spaces. This meticulous optimization process ensures that points close in the original space remain close in the reduced space, and points far apart remain distant. The algorithm's strength lies in its ability to reveal tightly packed clusters, making it exceptionally valuable for discerning subtle differences that define distinct biological entities or functional states within your sequence data. We activate this power to identify potential sub-types or evolutionary relationships that might be obscured in raw data.

Precision in parameter selection is paramount for t-SNE. The perplexity parameter acts as a control for the balance between local and global aspects of the data, effectively defining the number of nearest neighbors t-SNE considers. A common recommendation ranges between 5 and 50; selecting an appropriate value requires iterative testing and domain expertise. Lower perplexity focuses on the immediate neighborhood, while higher perplexity considers a broader context. Another critical parameter is iterations (max_iter), which dictates the number of optimization steps. Insufficient iterations can lead to incomplete convergence and a suboptimal projection. A typical starting point is 1000 iterations, with adjustment based on convergence behavior. While powerful for local structure, t-SNE presents challenges: its computational cost escalates quadratically with the number of data points, making it slower for massive datasets, and its non-deterministic nature means slight variations in output across runs, necessitating careful interpretation. Despite these, t-SNE provides unparalleled visual clarity for discovering local biological patterns.

# --- Apply t-SNE to Sequence Embeddings --- 
# t-SNE (t-Distributed Stochastic Neighbor Embedding) excels at revealing 
# local clusters and intricate groupings within the high-dimensional data. 
# We meticulously select parameters to optimize the visual output for biological interpretation.

# Define key t-SNE parameters:
# perplexity: Balances attention between local and global aspects of data. 
# A typical range is 5-50. Lower values focus on local structure, higher on global.
# iterations: Number of optimization steps. More iterations can lead to better convergence.
# check_duplicates: Set to FALSE for speed if you are certain there are no exact duplicate rows.

# Execute t-SNE algorithm
set.seed(456) # Ensure reproducibility for this t-SNE run
tsne_results <- Rtsne(
  sequence_embeddings_matrix, 
  dims = 2, 
  perplexity = 30, 
  theta = 0.5, 
  pca = FALSE, # Set to TRUE for faster computation with large datasets via initial PCA reduction
  max_iter = 1000,
  verbose = TRUE,
  is_distance = FALSE # Input is data matrix, not a distance matrix
)

# Extract the 2D coordinates from the t-SNE output
tsne_coords <- as.data.frame(tsne_results$Y)
colnames(tsne_coords) <- c("tSNE1", "tSNE2")

# Integrate t-SNE coordinates with original labels for visualization
tsne_plot_data <- cbind(tsne_coords, label = embeddings_df$label)

# Visualize t-SNE results using ggplot2
# We engineer a scatter plot to powerfully display distinct biological clusters.
plot_tsne <- ggplot(tsne_plot_data, aes(x = tSNE1, y = tSNE2, color = label)) +
  geom_point(alpha = 0.7, size = 2) +
  labs(
    title = "t-SNE Visualization of Sequence Embeddings",
    x = "t-SNE Dimension 1",
    y = "t-SNE Dimension 2",
    color = "Sequence Type"
  ) +
  theme_minimal() +
  theme(legend.position = "bottom", plot.title = element_text(hjust = 0.5, face = "bold"))

# Print the t-SNE plot to the R console (or save to file for persistent analysis)
print(plot_tsne)
Engineer UMAP for Global Context: Scalability and Manifold Exploration

Engineer UMAP for Global Context: Scalability and Manifold Exploration

UMAP emerges as a powerful, scalable dimensionality reduction technique, meticulously engineered to preserve both local and global structure within high-dimensional biological sequence embeddings. Unlike t-SNE’s probability-based approach, UMAP constructs a weighted graph in the high-dimensional space, representing the data’s manifold structure, and then optimizes a low-dimensional graph to be as structurally similar as possible. This graph-theoretic foundation confers distinct advantages: significantly faster computation, especially for large datasets, and a more consistent preservation of global relationships, ensuring that large-scale patterns and gradients are accurately reflected in the reduced space. For massive biological datasets, UMAP’s efficiency makes it an indispensable tool, allowing rapid exploration of vast genomic or proteomic landscapes.

Optimizing UMAP’s output necessitates a surgical approach to its key parameters. The n_neighbors parameter influences the balance between local and global structure. Smaller values force UMAP to focus on immediate neighborhoods, potentially yielding highly clustered local structures. Larger values encourage a more global view, reflecting the overall topological arrangement of the data. A robust starting point is often between 10 and 30, requiring iterative refinement based on the specific biological question. The min_dist parameter controls how tightly points are allowed to pack together in the low-dimensional representation. Lower values (e.g., 0.001) allow very dense clusters, while higher values (e.g., 0.5) distribute points more evenly. The choice of metric is also crucial; while Euclidean distance is a default, for sequence embeddings derived from neural networks, a cosine similarity metric often aligns better with biological interpretability, as it captures directional similarity irrespective of vector magnitude. By judiciously activating and tuning these parameters, we empower UMAP to decode the intricate manifold structure of our biological sequence embeddings, delivering clear and actionable visual insights.

# --- Apply UMAP to Sequence Embeddings --- 
# UMAP (Uniform Manifold Approximation and Projection) offers a compelling 
# alternative to t-SNE, excelling in preserving both local and global data structure 
# with superior computational efficiency, especially vital for large biological datasets.

# Define crucial UMAP parameters:
# n_neighbors: Controls the balance between local and global structure preservation. 
# Smaller values emphasize local structure, larger values emphasize global structure. 
# A common range is 5-50.
# min_dist: Governs how tightly points are packed together. 
# Lower values lead to denser clusters, higher values to more spread-out, balanced representations.
# metric: Defines the distance metric (e.g., "euclidean", "cosine"). 
# For embeddings, cosine similarity is often highly effective.

# Execute UMAP algorithm
set.seed(789) # Ensure reproducibility for this UMAP run
umap_results <- umap(
  sequence_embeddings_matrix,
  n_components = 2, 
  n_neighbors = 15, 
  min_dist = 0.1, 
  metric = "euclidean" # For sequence embeddings, 'cosine' is often more biologically relevant
)

# Extract the 2D coordinates from the UMAP output
umap_coords <- as.data.frame(umap_results$layout)
colnames(umap_coords) <- c("UMAP1", "UMAP2")

# Integrate UMAP coordinates with original labels for visualization
umap_plot_data <- cbind(umap_coords, label = embeddings_df$label)

# Visualize UMAP results using ggplot2
# We engineer a scatter plot to powerfully display distinct biological clusters,
# often with clearer separation of global structures than t-SNE.
plot_umap <- ggplot(umap_plot_data, aes(x = UMAP1, y = UMAP2, color = label)) +
  geom_point(alpha = 0.7, size = 2) +
  labs(
    title = "UMAP Visualization of Sequence Embeddings",
    x = "UMAP Dimension 1",
    y = "UMAP Dimension 2",
    color = "Sequence Type"
  ) +
  theme_minimal() +
  theme(legend.position = "bottom", plot.title = element_text(hjust = 0.5, face = "bold"))

# Print the UMAP plot
print(plot_umap)
Forging Visual Narratives: Comparative Analysis and Actionable Insights in R

Forging Visual Narratives: Comparative Analysis and Actionable Insights in R

Forging compelling visual narratives from our t-SNE and UMAP outputs transforms raw data into actionable biological insights. The true power emerges not just from generating plots, but from a strategic comparative analysis of both techniques and the judicious integration of external metadata. t-SNE typically excels at exposing granular, local relationships, often presenting distinct, well-separated clusters, even if the overall global layout appears somewhat arbitrary. Conversely, UMAP generally provides a more faithful representation of the data’s global topology, making it superior for visualizing continuous trajectories or hierarchical relationships between larger groups. By examining both projections side-by-side, we gain a comprehensive understanding: t-SNE confirms the tight cohesion of local biological groups, while UMAP validates their relative positioning within the broader biological landscape. We use ggplot2 in R to engineer high-quality, customizable visualizations that illuminate these distinctions.

Interpreting these projected landscapes requires a proactive, investigative mindset. First, identify distinct clusters: are there clear groupings of sequences? These often correspond to known protein families, gene isoforms, or species-specific sequences. Second, observe gradients or trajectories: do points transition smoothly between clusters? This might indicate evolutionary pathways, developmental stages, or functional spectrums. Third, pinpoint outliers: isolated points or small, distant clusters warrant immediate investigation, potentially revealing novel sequences, contamination, or unique functional variants. Crucially, integrate metadata. By mapping biological annotations – such as organism, known function, experimental condition, or evolutionary origin – onto the t-SNE and UMAP plots, we validate our observed clusters and imbue them with biological meaning. A cluster of sequences colored by 'Enzyme Family X' that also groups perfectly by 'Catalytic Activity A' provides robust, actionable evidence. This layered visualization activates a deeper level of biological understanding, transforming abstract coordinates into a vibrant, interpretable map of biological diversity and function.

# --- Comparative Visualization and Integration of Metadata --- 
# We combine the results of t-SNE and UMAP, leveraging ggplot2's capabilities 
# to construct insightful visual narratives from our sequence embeddings. 
# This step empowers us to compare algorithmic strengths and integrate critical metadata.

# For robust plotting, combine both t-SNE and UMAP results with original labels.
# Create a common data frame for plotting
combined_plot_data <- data.frame(
  tSNE1 = tsne_coords$tSNE1,
  tSNE2 = tsne_coords$tSNE2,
  UMAP1 = umap_coords$UMAP1,
  UMAP2 = umap_coords$UMAP2,
  label = embeddings_df$label
)

# Introduce synthetic metadata for demonstration (e.g., 'Source_Organism', 'Experimental_Condition')
# In a real pipeline, this would be loaded from an external file or database.
set.seed(999)
combined_plot_data$Source_Organism <- sample(c("E.coli", "S.cerevisiae", "H.sapiens"), 
                                             num_sequences, replace = TRUE,
                                             prob = c(0.4, 0.3, 0.3))
combined_plot_data$Experimental_Condition <- sample(c("Treatment", "Control"), 
                                                    num_sequences, replace = TRUE)

# --- Enhanced t-SNE Plot with Metadata --- 
plot_tsne_enhanced <- ggplot(combined_plot_data, aes(x = tSNE1, y = tSNE2, color = label, shape = Source_Organism)) +
  geom_point(alpha = 0.7, size = 3) +
  labs(
    title = "t-SNE of Embeddings: Clusters by Type, Shapes by Organism",
    x = "t-SNE Dimension 1",
    y = "t-SNE Dimension 2",
    color = "Sequence Type",
    shape = "Organism"
  ) +
  theme_bw() +
  theme(legend.position = "right", plot.title = element_text(hjust = 0.5, face = "bold"))

print(plot_tsne_enhanced)

# --- Enhanced UMAP Plot with Metadata --- 
plot_umap_enhanced <- ggplot(combined_plot_data, aes(x = UMAP1, y = UMAP2, color = label, shape = Experimental_Condition)) +
  geom_point(alpha = 0.7, size = 3) +
  labs(
    title = "UMAP of Embeddings: Clusters by Type, Shapes by Condition",
    x = "UMAP Dimension 1",
    y = "UMAP Dimension 2",
    color = "Sequence Type",
    shape = "Condition"
  ) +
  theme_bw() +
  theme(legend.position = "right", plot.title = element_text(hjust = 0.5, face = "bold"))

print(plot_umap_enhanced)

# --- Best Practices for Interpretation --- 
# 1. Look for distinct clusters: Indicates biologically meaningful groups.
# 2. Identify gradients/trajectories: Suggests continuous biological processes.
# 3. Assess outlier points: Potential novel sequences, contamination, or errors.
# 4. Integrate metadata: Overlaying biological annotations (e.g., function, species) 
#    validates clusters and provides deeper context.
# 5. Compare t-SNE and UMAP: t-SNE for local detail, UMAP for global structure. 
#    Discordances highlight areas for further investigation.
Optimize Embeddings: Navigate Pitfalls and Maximize Interpretative Leverage

Optimize Embeddings: Navigate Pitfalls and Maximize Interpretative Leverage

Optimizing the application of t-SNE and UMAP for sequence embeddings extends beyond basic execution; it demands a strategic approach to data pre-processing and a vigilant awareness of common pitfalls. For particularly high-dimensional or expansive datasets, an initial Principal Component Analysis (PCA) can serve as a potent pre-reduction step. While PCA is linear, it effectively de-correlates features and concentrates the majority of variance into a smaller number of principal components, significantly reducing the input dimensionality for t-SNE or UMAP. This not only accelerates computation but also helps filter out noise, leading to clearer, more robust projections. We activate PCA as a preparatory stage, ensuring our embeddings are optimally conditioned for subsequent non-linear reduction. The goal is to retain sufficient variance (e.g., 90-95%) in the PCA-reduced space before applying t-SNE or UMAP, thereby maximizing interpretative leverage without compromising computational efficiency.

Navigating the landscape of dimensionality reduction requires careful consideration of potential misinterpretations and adherence to best practices. A critical pitfall is misinterpreting distances: the exact distances between points in t-SNE or UMAP plots do not directly correspond to Euclidean distances in the original high-dimensional space. Focus instead on the presence of clusters, their relative proximity, and the overall topology. Another common error is over-interpreting small, isolated clusters, which might represent noise or artifacts rather than significant biological signals; always validate these with domain knowledge or additional statistical analysis. Neglecting parameter tuning is also detrimental; the default parameters for perplexity (t-SNE) and n_neighbors/min_dist (UMAP) are rarely optimal for the unique characteristics of biological sequence embeddings. Iterative testing and visualization across a range of parameters are crucial for robust results. Finally, always integrate external validation: the most compelling biological insights emerge when visual clusters are confirmed by independent biological metadata, functional assays, or statistical tests. By diligently applying these optimization strategies and avoiding common errors, we empower ourselves to decode maximum interpretative leverage from our sequence embeddings.

# --- Pre-processing Embeddings with PCA for Optimization --- 
# For exceptionally high-dimensional embeddings or very large datasets, 
# an initial Principal Component Analysis (PCA) can optimize performance 
# by reducing noise and further compressing dimensions before applying t-SNE or UMAP.
# This strategy preserves variance while accelerating downstream computations.

# First, center and scale the embedding matrix, a best practice for PCA.
scaled_embeddings <- scale(sequence_embeddings_matrix, center = TRUE, scale = TRUE)

# Execute PCA
pca_results <- prcomp(scaled_embeddings)

# Decide on the number of components to retain. 
# We typically aim to retain 90-95% of the variance or select components that capture 
# meaningful biological signal. For demonstration, we select 50 components.
num_pca_components <- 50
pca_reduced_embeddings <- pca_results$x[, 1:num_pca_components]

# Validate the dimensions of the PCA-reduced embeddings
print("Dimensions of PCA-reduced embeddings:")
print(dim(pca_reduced_embeddings))

# Now, apply t-SNE or UMAP to the PCA-reduced embeddings.
# For example, re-running UMAP on PCA-reduced data:
set.seed(111) # Reproducibility
umap_pca_results <- umap(
  pca_reduced_embeddings,
  n_components = 2, 
  n_neighbors = 15, 
  min_dist = 0.1, 
  metric = "euclidean" # Or 'cosine' if preferred for embeddings
)

umap_pca_coords <- as.data.frame(umap_pca_results$layout)
colnames(umap_pca_coords) <- c("UMAP_PCA1", "UMAP_PCA2")

# Visualize UMAP on PCA-reduced data
plot_umap_pca <- ggplot(cbind(umap_pca_coords, label = embeddings_df$label),
                       aes(x = UMAP_PCA1, y = UMAP_PCA2, color = label)) +
  geom_point(alpha = 0.7, size = 2) +
  labs(
    title = "UMAP on PCA-Reduced Embeddings",
    x = "UMAP Dimension 1 (from PCA)",
    y = "UMAP Dimension 2 (from PCA)",
    color = "Sequence Type"
  ) +
  theme_minimal() +
  theme(legend.position = "bottom", plot.title = element_text(hjust = 0.5, face = "bold"))

print(plot_umap_pca)

# --- Common Pitfalls and Best Practices --- 
# 1. Misinterpreting distances: Distances in t-SNE/UMAP plots are not always Euclidean in the original space.
#    Focus on clusters and relative positioning, not precise distances.
# 2. Over-interpreting small clusters: Small, isolated groups might be noise. Validate with biological knowledge.
# 3. Ignoring parameter tuning: Default parameters are rarely optimal for complex biological data. 
#    Iteratively tune perplexity, n_neighbors, and min_dist.
# 4. Over-reliance on 2D: Consider 3D visualization (e.g., with 'rgl' package) for complex structures.
# 5. Lack of validation: Always validate observed clusters with external biological data or statistical tests.
# 6. Pre-processing neglect: Ensure embeddings are properly normalized or scaled before DR.

Key Takeaways

Unveiling High-Dimensional Complexity

Biological sequence embeddings compress vast information into complex, high-dimensional vectors. Direct visualization is impossible, necessitating dimensionality reduction techniques like t-SNE and UMAP to uncover hidden structures and relationships within your biological data.

Mastering t-SNE for Local Precision

t-SNE excels at preserving local neighborhood structures, creating distinct, tight clusters that highlight fine-grained differences. Critical parameters like perplexity (5-50) and max_iter (e.g., 1000) require careful tuning. It is powerful for local detail but computationally intensive and non-deterministic.

Engineering UMAP for Global Perspective

UMAP offers superior computational efficiency and preserves both local and global data topology. Key parameters like n_neighbors (5-50) and min_dist (0.001-0.5) control the balance between these structures. It is often preferred for large datasets and visualizing broad trends or trajectories.

Forging Actionable Visual Narratives

Compare t-SNE (local details) and UMAP (global structure) to gain comprehensive insights. Integrate external biological metadata (e.g., species, function) into your plots using ggplot2 to validate clusters and transform abstract coordinates into meaningful biological interpretations. Look for clusters, gradients, and outliers.

Optimizing & Navigating Pitfalls

Pre-process high-dimensional embeddings with PCA to reduce noise and accelerate t-SNE/UMAP. Avoid misinterpreting distances, over-interpreting small clusters, and neglecting parameter tuning. Always validate visual findings with biological knowledge and additional statistical analysis for robust conclusions.

FAQ

  • When should I choose t-SNE over UMAP for my sequence embeddings?

    Choose t-SNE when your primary objective is to reveal fine-grained, local clustering structures within your sequence embeddings. t-SNE excels at separating distinct, tightly packed groups, even if the overall global arrangement of these groups seems less intuitive. If computational speed is not a primary concern and your dataset size is moderate (typically <100,000 points), t-SNE can provide unparalleled clarity for discerning subtle local differences between biological sequences.
  • When is UMAP the preferred method for visualizing sequence embeddings?

    UMAP is the preferred method when you require faster computation, particularly for large biological datasets (>100,000 sequences), and when preserving the global topological structure alongside local relationships is crucial. UMAP often yields a more consistent and interpretable global layout, making it excellent for visualizing continuous biological trajectories, hierarchical relationships, or large-scale patterns across diverse sequence families. Its deterministic nature also ensures reproducible results across multiple runs.
  • How do I determine optimal parameters for t-SNE (perplexity) and UMAP (n_neighbors, min_dist) for my sequence data?

    Determining optimal parameters is an iterative process requiring experimentation and biological context. For t-SNE's perplexity, test values typically between 5 and 50. For UMAP, try n_neighbors between 5 and 50 and min_dist between 0.001 and 0.5. Plot the results for various parameter combinations and evaluate which projection best reflects expected biological groupings or reveals new, interpretable patterns. Integrating metadata is key to validating your choices. There is no single 'best' parameter set; it's highly data- and question-dependent.
  • Can I use t-SNE or UMAP on sequence embeddings that have thousands of dimensions?

    Yes, both t-SNE and UMAP are designed precisely for high-dimensional data, including sequence embeddings with thousands of dimensions. However, for extremely high dimensions, it's a strong best practice to perform an initial dimensionality reduction step using PCA (Principal Component Analysis). This pre-processing step can effectively reduce noise and condense the data into a more manageable number of components (e.g., 50-300 principal components) before applying t-SNE or UMAP, leading to clearer results and faster computation.