Validate Protein Transformer Embeddings Accuracy in Python

Validate Protein Transformer Embeddings Accuracy in Python

The frontier of protein science accelerates with the advent of sophisticated artificial intelligence. We confront a pivotal challenge: translating complex biological sequences into actionable insights. Protein language models, particularly transformer architectures, redefine our ability to encode protein structure, function, and evolutionary relationships into high-dimensional numerical representations – embeddings. These embeddings promise unprecedented predictive power for tasks from drug discovery to enzyme engineering. Yet, their true utility hinges on a critical factor: accuracy. How reliably do these vectorized representations capture the biological reality of a protein?


This article empowers you to rigorously evaluate the performance of these transformative tools. We dissect the methodologies, benchmark datasets, and Pythonic scripts essential for asserting the fidelity of transformer embeddings, thereby enabling us to leverage the full potential of AI-driven protein analysis. Prepare to forge a robust framework for assessing embedding quality, ensuring your bio-engineering pipelines operate with unparalleled precision.

Activate the Embedded Revolution: Fundamentals and Setup

Activate the Embedded Revolution: Fundamentals and Setup

We initiate our journey into protein transformer embeddings by establishing a foundational understanding and a robust computational environment. Protein transformer embeddings represent a paradigm shift, transforming linear amino acid sequences into dense, context-aware numerical vectors. These vectors encapsulate intricate biological information, spanning evolutionary conservation, structural motifs, and functional properties. Their power originates from the transformer's self-attention mechanism, which meticulously weighs the importance of each amino acid in relation to all others within a sequence, thus capturing long-range dependencies and global context – a critical advancement over previous methods like Word2Vec or simple recurrent neural networks.


The accuracy of these embeddings directly impacts the reliability of any downstream machine learning task, from predicting protein-protein interactions to designing novel enzymes. A robust evaluation framework becomes indispensable. We advocate for a systematic approach, starting with precise environment setup in Python. Our standard practice mandates using virtual environments (e.g., conda or venv) to isolate project dependencies, preventing version conflicts and ensuring reproducibility. We specify core libraries: transformers for accessing pre-trained models, torch (or TensorFlow) as the deep learning backend, pandas and numpy for data manipulation, scikit-learn for downstream model training and evaluation, and biopython for sequence handling. This setup establishes the bedrock for all subsequent analytical operations, empowering us to seamlessly load state-of-the-art protein language models and embark on the rigorous evaluation process.

# Python Environment Setup
# We recommend creating a dedicated virtual environment for dependency management.
# For example, using conda:
# conda create -n protein_env python=3.9
# conda activate protein_env

# Install essential libraries
# We mandate specific versions to ensure reproducibility and compatibility.
pip install transformers==4.35.2 torch==2.1.0 pandas==2.1.4 scikit-learn==1.3.2 biopython==1.83 numpy==1.26.2

# Verify installations
import transformers
import torch
import pandas as pd
import sklearn
import Bio
import numpy as np

print(f"Transformers version: {transformers.__version__}")
print(f"PyTorch version: {torch.__version__}")
print(f"Pandas version: {pd.__version__}")
print(f"Scikit-learn version: {sklearn.__version__}")
print(f"Biopython version: {Bio.__version__}")
print(f"Numpy version: {np.__version__}")

# Dummy code to load an ESM model and extract embeddings for a single sequence
# We employ this for an initial sanity check of our setup.
from transformers import EsmTokenizer, EsmModel

def get_esm_embedding(sequence: str, model_name: str = "facebook/esm2_t6_8M_UR50D"):
    """
    Extracts ESM transformer embeddings for a given protein sequence.
    """
    tokenizer = EsmTokenizer.from_pretrained(model_name)
    model = EsmModel.from_pretrained(model_name)
    
    # Prepare sequence for model input
    # We prepend and append special tokens as required by ESM models.
    # See: https://github.com/facebookresearch/esm#pretrained-models
    data = [("protein1", sequence)]
    batch_converter = tokenizer.get_batch_converter()
    batch_labels, batch_strs, batch_tokens = batch_converter(data)
    
    with torch.no_grad():
        results = model(batch_tokens, repr_layers=[6], return_dict=True)
    
    # Extract per-sequence average embedding from the last hidden state (or a specified layer)
    # We average over all non-special tokens for a concise representation.
    token_embeddings = results.last_hidden_state[0, 1:-1] # Remove CLS and EOS tokens
    sequence_embedding = token_embeddings.mean(dim=0).numpy()
    
    return sequence_embedding

# Example usage:
protein_sequence = "MGLSDGEWQLVLNVWGKVEADIPGHGQEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG"
embedding = get_esm_embedding(protein_sequence)
print(f"\nExtracted embedding shape: {embedding.shape}")
print(f"First 5 values of embedding: {embedding[:5]}")

Engineer Robust Evaluation: Datasets and Metrics Selection

We move to the critical phase of engineering a robust evaluation strategy, commencing with the meticulous selection of benchmark datasets and appropriate performance metrics. The fidelity of protein transformer embeddings is contingent upon their ability to generalize across diverse biological tasks. We therefore activate a portfolio of benchmark datasets that rigorously test their capabilities. For classification tasks, such as predicting protein subcellular localization or secondary structure, datasets like those compiled within the TAPE benchmark or curated subsets of UniProt are invaluable. For regression tasks, which include predicting binding affinity, thermal stability, or enzyme kinetics, specialized experimental datasets are required. We meticulously ensure dataset balance to prevent model bias and employ stratified sampling during data splitting to maintain class proportions across training and testing sets.


The choice of evaluation metrics fundamentally shapes our understanding of model performance. For binary and multi-class classification, we prioritize metrics beyond simple accuracy, which can be misleading on imbalanced datasets. We deploy Precision, Recall, F1-score, and the Matthews Correlation Coefficient (MCC), which offers a balanced measure even with class imbalances. For probabilistic models, the Area Under the Receiver Operating Characteristic curve (AUC-ROC) provides a robust assessment of classifier discrimination. In regression scenarios, we leverage Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Spearman's or Pearson's correlation coefficient to quantify the strength and direction of linear and monotonic relationships. Employing cross-validation techniques (e.g., k-fold) is non-negotiable; it guarantees a more reliable estimate of generalization performance, mitigating the risks of overfitting to a single train-test split. This disciplined approach secures a comprehensive and unbiased assessment of embedding accuracy.

# Placeholder for benchmark dataset loading and preprocessing logic
# We illustrate how to structure data for a classification task.

import pandas as pd
from sklearn.model_selection import train_test_split

# Assume 'dataset.csv' contains protein sequences and their labels
# Example structure: 'sequence', 'label'
# For demonstration, we create a dummy dataset.

data = {
    'sequence': [
        'MGLSDGEWQLVLNVWGKVEADIPGHGQEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG',
        'MSILGYWKLVEKVGKVVDSKPGHGQDLILKLFKQHPETIEKFDRIKYLKSKDDEMKASADLKKHG',
        'MAAASSDSKLSVPMKVVVATVGVTGAYGSDVEQLIRKAFEPHPEELIKFFDKLKSKSLESMKASEDLKKHG',
        'MVVKALWKLVKKVGKVVDPKPGHGQDIIIKLFEQHPETLEKFDRVKHLKSDDKLMKASADLKKHG',
        'MSILEYWKLVSAVGKVVDSSPGHGQELILKLFKQHPETLEKFDRVKYLKTKDDEMKASADLKKHG',
        'MMGLEYWQLVLAVGKVEADIPGHGQEVIIRLFKGHPETLEKFDKFKHLKSEDEMKASADLKKHG',
        'MLAKLGYWQLVLNVWGKVEADIPGHGHEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG',
        'MAASSGEWQLVLNVWGKVEADIPGHGQEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG'
    ],
    'label': [
        0, 1, 0, 1, 0, 1, 0, 1
    ]
}
df = pd.DataFrame(data)

# We split the data into training and testing sets to simulate real-world evaluation.
X_train, X_test, y_train, y_test = train_test_split(df['sequence'], df['label'], test_size=0.25, random_state=42)

print(f"Training sequences count: {len(X_train)}")
print(f"Testing sequences count: {len(X_test)}")
print(f"Training labels distribution:\n{y_train.value_counts()}")
print(f"Testing labels distribution:\n{y_test.value_counts()}")

# Example of how to structure evaluation metrics for a classification task
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, matthews_corrcoef, roc_auc_score

def evaluate_classifier(y_true, y_pred, y_prob=None):
    """
    Calculates and prints common classification metrics.
    """
    print(f"Accuracy: {accuracy_score(y_true, y_pred):.4f}")
    print(f"Precision (macro): {precision_score(y_true, y_pred, average='macro'):.4f}")
    print(f"Recall (macro): {recall_score(y_true, y_pred, average='macro'):.4f}")
    print(f"F1-Score (macro): {f1_score(y_true, y_pred, average='macro'):.4f}")
    print(f"MCC: {matthews_corrcoef(y_true, y_pred):.4f}")
    if y_prob is not None and len(np.unique(y_true)) > 1:
        print(f"AUC-ROC: {roc_auc_score(y_true, y_prob):.4f}")

# Dummy predictions for demonstration
dummy_y_true = np.array([0, 1, 0, 1, 0, 1, 0, 1])
dummy_y_pred = np.array([0, 0, 0, 1, 1, 1, 0, 1])
dummy_y_prob = np.array([0.1, 0.2, 0.3, 0.8, 0.7, 0.9, 0.4, 0.6])

print("\n--- Dummy Classification Metrics --- ")
evaluate_classifier(dummy_y_true, dummy_y_pred, dummy_y_prob)
Decode Performance in Python: Implementation Workflow

Decode Performance in Python: Implementation Workflow

We now execute the core implementation workflow in Python, decoding the performance of protein transformer embeddings through a systematic pipeline. This process involves four critical steps: loading the pre-trained transformer model, extracting embeddings from our protein sequences, training a lightweight downstream model on these embeddings, and finally, evaluating its performance. We primarily leverage models from the Hugging Face transformers library, which provides seamless access to models like ESM (Evolutionary Scale Modeling) or ProtT5. Our methodology ensures that the pre-trained model is set to evaluation mode (model.eval()) to disable dropout layers and batch normalization updates, guaranteeing consistent inference.


For embedding extraction, we tokenize each protein sequence according to the model's specific tokenizer (e.g., adding <CLS> and <EOS> tokens for ESM models) and then pass the tokenized input through the model. We typically extract the mean-pooled last hidden state of the transformer layers. This produces a fixed-size vector representing the entire protein sequence, suitable for downstream machine learning tasks. Crucially, we perform this operation within a torch.no_grad() context to conserve memory and optimize computation, as gradient calculations are unnecessary for inference. Subsequently, these extracted embeddings serve as features for a simple, fast-training downstream classifier or regressor, such as LogisticRegression, RandomForestClassifier, or Support Vector Machine from scikit-learn. The choice of a lightweight model isolates the performance contribution of the embeddings themselves, rather than obfuscating it with complex downstream model architectures. We apply k-fold cross-validation during this training phase, ensuring our performance metrics are robust and indicative of true generalization capability. This rigorous implementation strategy unveils the intrinsic accuracy potential of the transformer embeddings.

# Full Python script for extracting embeddings and training a downstream classifier
# We ensure each step is transparent and reproducible.

import torch
from transformers import EsmTokenizer, EsmModel
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split, StratifiedKFold
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score, matthews_corrcoef, roc_auc_score

# 1. Define the protein sequences and labels (example data)
# We construct a synthetic dataset for illustrative purposes.
sequences = [
    "MGLSDGEWQLVLNVWGKVEADIPGHGQEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG",
    "MSILGYWKLVEKVGKVVDSKPGHGQDLILKLFKQHPETIEKFDRIKYLKSKDDEMKASADLKKHG",
    "MAAASSDSKLSVPMKVVVATVGVTGAYGSDVEQLIRKAFEPHPEELIKFFDKLKSKSLESMKASEDLKKHG",
    "MVVKALWKLVKKVGKVVDPKPGHGQDIIIKLFEQHPETLEKFDRVKHLKSDDKLMKASADLKKHG",
    "MSILEYWKLVSAVGKVVDSSPGHGQELILKLFKQHPETLEKFDRVKYLKTKDDEMKASADLKKHG",
    "MMGLEYWQLVLAVGKVEADIPGHGQEVIIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG",
    "MLAKLGYWQLVLNVWGKVEADIPGHGHEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG",
    "MAASSGEWQLVLNVWGKVEADIPGHGQEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG",
    "MTTFLKLVKNWGKVVNNIPGHGQEVLISLFKGHPDTLERFDKFTHLKSEDEMKASEDLKKHG",
    "MMGLSDGEWQLVLNVWGKVEADIPGHGQEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG",
    "MSEELGYWKLVEKVGKVVDSKPGHGQDLILKLFKQHPETIEKFDRIKYLKSKDDEMKASADLKKHG",
    "MAAASSDSKLSVPMKVVVATVGVTGAYGSDVEQLIRKAFEPHPEELIKFFDKLKSKSLESMKASEDLKKHG",
    "MVVKALWKLVKKVGKVVDPKPGHGQDIIIKLFEQHPETLEKFDRVKHLKSDDKLMKASADLKKHG",
    "MSILEYWKLVSAVGKVVDSSPGHGQELILKLFKQHPETLEKFDRVKYLKTKDDEMKASADLKKHG",
    "MMGLEYWQLVLAVGKVEADIPGHGQEVIIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG",
    "MLAKLGYWQLVLNVWGKVEADIPGHGHEVLIRLFKGHPETLEKFDKFKHLKSEDEMKASEDLKKHG"
]
labels = [
    0, 1, 0, 1, 0, 1, 0, 1,
    0, 1, 0, 1, 0, 1, 0, 1
]

# 2. Load pre-trained ESM model and tokenizer
# We opt for a smaller ESM-2 model (esm2_t6_8M_UR50D) for quicker execution.
model_name = "facebook/esm2_t6_8M_UR50D"
tokenizer = EsmTokenizer.from_pretrained(model_name)
model = EsmModel.from_pretrained(model_name)
model.eval() # Set model to evaluation mode

# Move model to GPU if available
if torch.cuda.is_available():
    model = model.to('cuda')
    print("Model moved to GPU.")
else:
    print("Model running on CPU.")

# 3. Function to extract embeddings
def get_sequence_embedding(sequence: str, tokenizer, model, device='cpu'):
    """
    Extracts mean-pooled embeddings for a single protein sequence.
    We handle tokenization, model inference, and pooling.
    """
    with torch.no_grad():
        # Prepare input. Add CLS and EOS tokens.
        token_ids = tokenizer.encode(sequence, add_special_tokens=True, return_tensors='pt').to(device)
        
        # Forward pass through the model
        output = model(token_ids, output_hidden_states=True)
        
        # Extract the last hidden state and remove special tokens (CLS, EOS)
        # We select the last hidden state for the most contextualized embeddings.
        embeddings = output.last_hidden_state[0, 1:-1] # shape: (seq_len, embedding_dim)
        
        # Mean-pool to get a fixed-size sequence embedding
        mean_embedding = embeddings.mean(dim=0).cpu().numpy()
    return mean_embedding

# 4. Extract embeddings for all sequences
all_embeddings = []
device = 'cuda' if torch.cuda.is_available() else 'cpu'
for seq in sequences:
    emb = get_sequence_embedding(seq, tokenizer, model, device=device)
    all_embeddings.append(emb)

X = np.array(all_embeddings)
y = np.array(labels)

print(f"\nShape of all embeddings: {X.shape}") # Should be (num_sequences, embedding_dim)

# 5. Downstream task: Train a Logistic Regression classifier with cross-validation
# We employ StratifiedKFold to maintain class distribution in each fold.
kf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

accuracy_scores = []
f1_scores = []
mcc_scores = []
auc_roc_scores = []

print("\n--- Starting K-Fold Cross-Validation --- ")
for fold, (train_index, test_index) in enumerate(kf.split(X, y)):
    X_train, X_test = X[train_index], X[test_index]
    y_train, y_test = y[train_index], y[test_index]
    
    # We select a simple Logistic Regression as a lightweight classifier.
    classifier = LogisticRegression(max_iter=1000, random_state=42, solver='liblinear')
    classifier.fit(X_train, y_train)
    y_pred = classifier.predict(X_test)
    y_prob = classifier.predict_proba(X_test)[:, 1] # Probability of the positive class
    
    accuracy_scores.append(accuracy_score(y_test, y_pred))
    f1_scores.append(f1_score(y_test, y_pred))
    mcc_scores.append(matthews_corrcoef(y_test, y_pred))
    
    # Only compute AUC-ROC if there's more than one class in y_test
    if len(np.unique(y_test)) > 1:
        auc_roc_scores.append(roc_auc_score(y_test, y_prob))
    else:
        auc_roc_scores.append(np.nan) # Assign NaN if AUC cannot be computed

    print(f"Fold {fold+1}: Accuracy={accuracy_scores[-1]:.4f}, F1-score={f1_scores[-1]:.4f}, MCC={mcc_scores[-1]:.4f}, AUC-ROC={auc_roc_scores[-1]:.4f}")

print("\n--- Aggregated Results --- ")
print(f"Average Accuracy: {np.mean(accuracy_scores):.4f} +/- {np.std(accuracy_scores):.4f}")
print(f"Average F1-score: {np.mean(f1_scores):.4f} +/- {np.std(f1_scores):.4f}")
print(f"Average MCC: {np.mean(mcc_scores):.4f} +/- {np.std(mcc_scores):.4f}")

# Filter out NaN values for AUC-ROC if any fold had only one class
valid_auc_roc_scores = [score for score in auc_roc_scores if not np.isnan(score)]
if valid_auc_roc_scores:
    print(f"Average AUC-ROC: {np.mean(valid_auc_roc_scores):.4f} +/- {np.std(valid_auc_roc_scores):.4f}")
else:
    print("AUC-ROC could not be computed for any fold due to single-class test sets.")

Optimize for Precision: Interpretation and Best Practices

We finalize our evaluation by meticulously interpreting the generated results and establishing best practices for optimizing embedding precision. The average and standard deviation of metrics across cross-validation folds provide a quantitative measure of embedding accuracy and robustness. A high average metric (e.g., F1-score, MCC, AUC-ROC) coupled with a low standard deviation indicates a stable and performant embedding representation. Conversely, significant variability suggests instability, potentially stemming from dataset characteristics, model sensitivity, or insufficient training data. We activate critical thinking: Is a 75% accuracy sufficient for your specific biological application, or does it mandate further refinement?


We confront common pitfalls head-on. Data leakage, where information from the test set inadvertently contaminates the training process, can inflate metrics misleadingly. Rigorous data splitting and feature engineering must prevent this. An over-parameterized downstream model might overfit the embeddings, making them appear more capable than they are; hence, we advocate for simpler models initially. To optimize precision, we consider several advanced strategies: fine-tuning the transformer model itself on your specific task's dataset, rather than solely using pre-trained embeddings, can dramatically enhance task-specific performance. Hyperparameter optimization for the downstream model (e.g., using GridSearchCV or RandomizedSearchCV) can squeeze out additional performance. Furthermore, ensemble methods combining predictions from multiple models or embeddings can yield more robust and accurate outcomes. We also emphasize the paramount importance of biological context: numerical accuracy must align with scientific plausibility. A high-performing model that contradicts known biological principles demands further investigation. We forge an iterative cycle of evaluation, interpretation, and refinement, pushing the boundaries of what protein transformer embeddings can achieve in your bio-engineering pipelines.

# No new code for this section. Focuses on discussion and interpretation of previous code's output.
# We assume the user has run the previous Python script and has the aggregated metrics.

# Example output from previous script for discussion (replace with actual run data):
# Average Accuracy: 0.7500 +/- 0.1000
# Average F1-score: 0.7300 +/- 0.1200
# Average MCC: 0.5000 +/- 0.2000
# Average AUC-ROC: 0.8000 +/- 0.1500

# Key discussion points for the user:
# 1. Analyze the mean and standard deviation of each metric across folds.
# 2. Compare F1, MCC, and AUC-ROC (if applicable) for a balanced view.
# 3. Consider dataset characteristics (imbalance, size, quality).
# 4. Discuss potential pitfalls like data leakage or overfitting of the downstream model.
# 5. Explore strategies for improvement (fine-tuning, hyperparameter optimization).
# 6. Emphasize biological relevance of the task and embedding performance.

Key Takeaways

Protein Transformer Embeddings: Foundational Concepts

Protein transformer embeddings translate amino acid sequences into rich numerical vectors, capturing structural and functional attributes through self-attention mechanisms. Their accuracy is paramount for reliable downstream biological applications. We mandate a disciplined Python environment setup using virtual environments and key libraries like transformers, torch, and scikit-learn, establishing the bedrock for robust evaluation.

Strategic Evaluation: Datasets and Metrics

Rigorous evaluation necessitates careful selection of benchmark datasets (e.g., TAPE, UniProt subsets) and appropriate metrics. For classification, we activate Precision, Recall, F1-score, MCC, and AUC-ROC. For regression, RMSE, MAE, and correlation coefficients are essential. Stratified k-fold cross-validation is non-negotiable to ensure unbiased, generalizable performance assessments and mitigate overfitting.

Pythonic Implementation for Performance Decoding

Our implementation workflow in Python involves loading pre-trained transformer models (e.g., ESM-2), meticulously extracting mean-pooled last hidden state embeddings for each protein, and then training a lightweight downstream classifier (e.g., Logistic Regression) on these features. We execute this process with torch.no_grad() for efficiency and employ k-fold cross-validation to provide robust performance metrics, directly decoding the embeddings' efficacy.

Optimizing for Precision: Interpretation and Best Practices

We interpret results by analyzing average metrics and standard deviations across folds, seeking high accuracy and low variability. We surgically address pitfalls like data leakage and downstream model overfitting. Optimization strategies include fine-tuning the transformer model, hyperparameter tuning of downstream models, and leveraging ensemble methods. We consistently cross-reference numerical accuracy with biological plausibility to ensure meaningful advancements.

FAQ

  • What is a protein transformer embedding?

    A protein transformer embedding is a dense numerical vector generated by a transformer-based neural network model. This vector captures the complex biochemical and biophysical properties of a protein sequence, including its evolutionary history, structural motifs, and functional characteristics, by analyzing amino acid dependencies and context.

  • Why is it important to evaluate the accuracy of these embeddings?

    Evaluating embedding accuracy is critical because the quality of these representations directly dictates the reliability and predictive power of any downstream machine learning task. Inaccurate embeddings can lead to erroneous biological insights, flawed drug designs, or inefficient protein engineering outcomes. Rigorous evaluation ensures the embeddings faithfully represent the underlying biological reality.

  • Which benchmark datasets are suitable for evaluating protein embeddings?

    Suitable benchmark datasets depend on the specific task. For general protein properties, datasets like those in the TAPE benchmark (e.g., secondary structure, remote homology detection) are excellent. For specialized tasks, curated datasets from resources like UniProt, experimentally validated protein-protein interaction datasets, or protein stability assays are invaluable. Always ensure data quality and relevance to your problem.

  • What are common pitfalls during protein embedding evaluation?

    Common pitfalls include data leakage (information from test set seeping into training), using inappropriate evaluation metrics for imbalanced datasets, overfitting of the downstream model, selecting a transformer model not suited for the task, and failing to perform adequate cross-validation. Neglecting biological context when interpreting results also represents a significant pitfall.

  • How can I improve the accuracy of my protein transformer embeddings?

    We recommend several strategies: fine-tuning the pre-trained transformer model on your specific, task-relevant dataset; judiciously selecting the optimal transformer layer for embedding extraction; applying hyperparameter optimization to your downstream model; exploring ensemble methods; and ensuring your training data is high-quality, diverse, and well-curated. Continuous iteration and validation against biological ground truth are key.