CHARM
CHARM is a zero-shot probabilistic time-series foundation model from C3 AI. It produces full quantile forecasts (99 quantile levels) for arbitrary horizons and is evaluated zero-shot on the GIFT-Eval benchmark β no GIFT-Eval data is used in training.
Availability: CHARM model weights and training/replication code are not publicly released. The model is served behind an API and consumed through the open-source
c3-charmPython SDK (see Using CHARM below). This page serves as the model card and benchmark record.
Model summary
| Parameters | ~63.3M (encoder ~59M) |
| Hidden size (d_model) | 384 |
| Projection | TCN-pool (patch size 16, causal conv stack + residual gate) |
| Backbone | Transformer encoder with RoPE attention |
| Decoder | Quantile decoder, 99 levels (0.01β0.99) |
| Max context length | 8192 |
| Output | Probabilistic (quantile) forecasts, multivariate-capable |
| Precision | float32 |
Intended use
Zero-shot probabilistic forecasting of univariate and multivariate time series across domains (energy, transport, sales, healthcare, nature, web/cloud-ops, econ/finance). The model is applied without any per-dataset fine-tuning.
Using CHARM β the c3-charm SDK
CHARM is served behind an API and consumed through the open-source c3-charm Python SDK. The SDK provides embeddings (multivariate time series β vectors), forecast/backcast (quantile predictions), and an optional toolkit for downstream tasks (anomaly detection, retrieval, classification, reconstruction, forecasting). The reference below is the full SDK documentation.
Dual-model serving. A CHARM server can be backed by two independent checkpoints β one for embeddings (
/predict,client.embeddings) and a separate one for forecasting (/forecast,client.prediction). They may differ in architecture, patch size, and embedding dimension, so read per-model properties fromclient.model_info()rather than hardcoding.
Example notebooks
Runnable notebooks live in notebooks/ in this repo. Open them on the Hub, or download the folder and run locally after pip install c3-charm[toolkit] (set CHARM_BASE_URL / CHARM_API_KEY first). The classification and reconstruction/forecasting demos read the small sample datasets in notebooks/data/.
| Notebook | What it covers |
|---|---|
getting_started.ipynb |
Client setup, first embeddings call, inspecting model_info(). |
charm_toolkit_demo.ipynb |
Tour of the charm_toolkit β datasets, precompute, trainer, and each task head. |
demo_forecasting.ipynb |
Zero-shot quantile forecasting and the embedding-based ForecastingModel head (weather data). |
demo_reconstruction.ipynb |
Backcast reconstruction and the ReconstructionModel head for anomaly detection (weather data). |
demo_classification.ipynb |
Time-series classification with the ClassificationModel head (BasicMotions data). |
demo_retrieval_anomaly_detection.ipynb |
Embedding retrieval + kNN / zero-shot anomaly-detection recipes. |
Installation
pip install c3-charm # core SDK only (embeddings + forecast)
pip install c3-charm[toolkit] # includes PyTorch models, datasets, trainers
Or from source:
git clone https://github.com/c3ai/c3-charm.git
cd c3-charm
poetry install # core SDK only
poetry install --with toolkit # include toolkit dependencies
CPU vs CUDA torch
pip install c3-charm pulls the default torch wheel from PyPI, which on
Linux is the full CUDA build (~4 GB). The SDK itself never requires a
GPU β the model runs server-side, and toolkit training heads are small
enough to fit on CPU β so if you want the smaller CPU-only wheel, install
torch from the CPU index before c3-charm:
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install c3-charm
Or from source, use bash setup.sh, which forces the CPU wheel.
If you instead want a specific CUDA build, install it manually after
c3-charm:
pip install --force-reinstall --no-deps torch \
--index-url https://download.pytorch.org/whl/cu124
Swap cu124 for the CUDA version that matches your driver (cu121,
cu126, cu128, β¦). --no-deps is important β it stops pip from
re-resolving your torch install.
Core SDK
Client initialization
from charm import CharmClient
client = CharmClient(
base_url="http://your-server:8080",
api_key="your-api-key", # or set CHARM_API_KEY env var
timeout=300, # override; SDK default is 15s β raise it for /forecast (server allows up to ~220s)
max_retries=3,
)
Embeddings β client.embeddings.create()
Converts time series windows into dense vectors.
response = client.embeddings.create(
descriptions=[["sensor_A", "sensor_B"]], # (N, C) channel names
ts_array=[[[1.0, 2.0], [1.1, 2.1], ...]], # (N, T, C) values
batch_size=32,
return_tensors="np", # "list", "np", or "torch"
aggregate=True, # True β (N, D); False β (N, T_, C, D)
progress=True,
)
embeddings = response.embeds # shape (N, D) when aggregate=True
aggregate parameter:
True(default): Returns flattened embeddings(N, D)β one vector per series. Best for retrieval, classification, clustering.False: Returns per-patch, per-channel embeddings(N, T_, C, D)whereT_ = ceil(T / patch_size). Best for fine-grained tasks or custom heads.
Async (faster for large datasets):
response = await client.embeddings.async_create(
descriptions=descriptions,
ts_array=ts_array,
max_B_per_request=32,
concurrency_per_call=8,
return_tensors="np",
aggregate=True,
)
Forecast / Backcast β client.prediction.create()
Zero-shot quantile predictions β no training required.
response = client.prediction.create(
descriptions=[["sensor_A", "sensor_B"]],
ts_array=[[[1.0, 2.0], [1.1, 2.1], ...]],
target_len=10, # positive = forecast, negative = backcast
return_tensors="np",
)
forecast = response.denormalized_predictions # (N, target_len=10, C, Q) β Q quantiles (model-dependent, e.g. 21)
median = response.median # (N, 10, C) β point forecast (median quantile)
Backcast (reconstruct past values):
response = client.prediction.create(
descriptions=descriptions,
ts_array=ts_array,
target_len=-8, # reconstruct last 8 steps
return_tensors="np",
)
Input constraints
| Constraint | Limit |
|---|---|
| Timesteps per series | 1 β€ T β€ 8192 (the model's training window) |
| Channels per series | No hard limit (large C grows memory ~O(CΒ²) under cross-channel attention) |
| Per-request size | N Γ C Γ T β€ 500,000 β SDK client-side batching guard, not a server limit |
| Batch consistency | All series in a request must share the same T and C |
| Best accuracy | T β€ 8192 (the model's training window); T need not be a multiple of the patch size |
The model was trained on windows of up to 8192 timesteps (patch size 16).
Inputs are not required to be a multiple of the patch size β they are padded
to a patch boundary internally. The architecture can technically accept more
than 8192 (up to 1500 patches), but longer inputs rely on positions seen only
in pretraining, so quality degrades. Query client.model_info() for the
served model's actual patch size and embedding dimension.
The SDK handles client-side batching automatically when you set batch_size (sync) or max_B_per_request (async).
Output shapes
| Method | Output field | Shape |
|---|---|---|
embeddings.create(aggregate=True) |
response.embeds |
(N, D) |
embeddings.create(aggregate=False) |
response.embeds |
(N, T_, C, D), T_ = ceil(T / patch_size) |
prediction.create(target_len > 0) |
response.denormalized_predictions |
(N, target_len, C, Q) |
prediction.create(target_len < 0) |
response.denormalized_predictions |
(N, abs(target_len), C, Q) |
prediction.create(...) |
response.median |
(N, abs(target_len), C) |
Channel descriptions
Descriptions are required and affect embedding quality. They tell the model what each channel represents.
Good descriptions β use meaningful, consistent names:
descriptions = [["engine_temperature", "oil_pressure", "rpm"]]
Acceptable β short but informative:
descriptions = [["temp", "pressure", "speed"]]
Avoid β generic or positional names reduce model effectiveness:
descriptions = [["col_0", "col_1", "col_2"]] # works but suboptimal
When working with pandas DataFrames, use column names directly:
descriptions = [df.columns.tolist()] * N
Scaling
No pre-processing needed. CHARM normalizes internally. Send raw data as-is. Do not apply StandardScaler, MinMaxScaler, or log transforms before calling the API.
Error handling
from charm import CharmError, AuthenticationError, InvalidRequestError, RateLimitError
try:
response = client.embeddings.create(...)
except AuthenticationError:
# bad API key
except InvalidRequestError as e:
# shape violations, empty input
except RateLimitError:
# back off and retry
except CharmError as e:
# catch-all for other SDK errors
Toolkit β Downstream Tasks
The toolkit (pip install c3-charm[toolkit]) provides PyTorch models, dataset utilities, and training infrastructure for fine-tuning on top of CHARM embeddings.
Retrieval β charm_toolkit.retrieval
Find similar time series by embedding similarity.
from charm_toolkit.retrieval import (
l2_normalize,
cosine_similarity_matrix,
knn_search,
retrieval_metrics,
)
# Embed your data
response = client.embeddings.create(
descriptions=descriptions,
ts_array=windows_list,
return_tensors="np",
)
embeddings = response.embeds # (N, D)
# Similarity search
sim = cosine_similarity_matrix(embeddings, embeddings)
# kNN search
indices, scores = knn_search(query_emb, corpus_emb, k=5)
# Evaluation metrics
metrics = retrieval_metrics(
query_emb=query_emb,
corpus_emb=corpus_emb,
query_labels=query_labels,
corpus_labels=corpus_labels,
k_values=[1, 3, 5, 10],
exclude_self=True,
query_ids=query_dataset_names,
corpus_ids=corpus_dataset_names,
)
# Returns: precision@k, ndcg@k, hit_rate@k
Anomaly Detection β charm_toolkit.anomaly_detection
Detect anomalies via kNN distance scoring on windowed CHARM embeddings.
from charm_toolkit.anomaly_detection import (
sliding_window_embeddings,
knn_anomaly_scores,
window_scores_to_pointwise,
)
# 1. Embed sliding windows
train_emb = sliding_window_embeddings(
client, train_data, descriptions,
window_size=128, stride=1, batch_size=64,
)
test_emb = sliding_window_embeddings(
client, test_data, descriptions,
window_size=128, stride=1, batch_size=64,
)
# 2. Score test windows by distance to train
window_scores = knn_anomaly_scores(
test_emb=test_emb,
reference_emb=train_emb,
k=5,
distance="cosine", # "cosine", "l2", "l1"
aggregation="mean", # "mean", "max"
)
# 3. Aggregate to per-timestep scores
pointwise_scores = window_scores_to_pointwise(
window_scores=window_scores,
window_size=128,
stride=1,
total_length=len(test_data),
method="mean", # "mean", "max", "last", "center"
)
Pointwise aggregation methods:
Each timestep is covered by multiple overlapping windows. The method parameter controls how to assign a single score per timestep:
| Method | Behavior | Use case |
|---|---|---|
"mean" |
Average of all windows covering the point | Smooth, best for offline evaluation |
"max" |
Max score among covering windows | Conservative, catches isolated spikes |
"last" |
Score of the most recently completed window | Online/streaming β score only updates when a window finishes processing |
"center" |
Score of the window centered on each point | Minimal time-shift, tightest temporal alignment |
Zero-shot recipes (no clean reference set required) β recommended methods, in order of strength:
- Bootstrap k-NN (best). Two steps: run
sklearn.ensemble.IsolationForeston the embedding matrix and take the bottom ~70% by score as a presumed-clean reference; then callknn_anomaly_scoresagainst that reference. - CBLOF:
pyod.models.cblof.CBLOFon the embedding matrix. - IsolationForest:
sklearn.ensemble.IsolationForestdirectly on the embedding matrix.
L2-normalize embeddings beforehand to use cosine geometry. See demo_retrieval_anomaly_detection.ipynb.
Best practices (validated on TSB-AD, metric VUS-PR):
Ensemble the embedding score with per-window statistics. The encoder instance-normalizes each window, erasing amplitude/level-shift anomalies (spikes, steps β the most common kind) from the embedding. Recover them with
window_statistics(no extra model call) and combine viaensemble_scores(method="zscore"β z-score each detector and sum; parameter-free and as strong as a tuned weighted ensemble). This is worth ~+7 pp VUS-PR supervised.from charm_toolkit.anomaly_detection import window_statistics, ensemble_scores emb_score = knn_anomaly_scores(test_emb, train_emb, k=3, distance="cosine") S_tr = window_statistics(train_data, 128, 1) mu, sd = S_tr.mean(0, keepdims=True), S_tr.std(0, keepdims=True) + 1e-8 stats_score = knn_anomaly_scores( (window_statistics(test_data, 128, 1) - mu) / sd, (S_tr - mu) / sd, k=3, distance="l2") window_scores = ensemble_scores([emb_score, stats_score], method="zscore") # zero-shot: build the reference by bootstrap (above) and use method="rank"Multivariate: do NOT pool channels when the channel count is high. Channel-mean pooling dilutes an anomaly confined to a few channels across all of them. Keep channels separate with
sliding_window_channel_embeddings+per_channel_knn_scores(channel_pool="adaptive"). The benefit grows with the number of channels β negligible at Cβ€3, sizeable at Cβ₯20, large at C>60 β so it is advisable whenever C is high. For few channels, plain mean-pooling is equally good and cheaper.from charm_toolkit.anomaly_detection import ( sliding_window_channel_embeddings, per_channel_knn_scores) train_pc = sliding_window_channel_embeddings(client, train_data, descriptions) # (N, C, D) test_pc = sliding_window_channel_embeddings(client, test_data, descriptions) emb_score = per_channel_knn_scores(test_pc, train_pc, k=3, channel_pool="adaptive")
ReconstructionModel β anomaly detection via learned head
from charm_toolkit import (
ReconstructionModel, create_reconstruction_datasets,
collator, TrainerClass,
)
from torch.utils.data import DataLoader
import torch.nn as nn
train_ds, val_ds, test_ds = create_reconstruction_datasets(
raw_data, # (T, C) numpy array or torch tensor
descriptions=channel_names,
window_size=256,
stride=1,
train_ratio=0.7,
val_ratio=0.15,
sequential=True,
scale=True,
)
model = ReconstructionModel(
embedding_client=client,
reconstructor="linear", # "linear", "mlp", or custom nn.Module
hidden_dim=128,
dropout=0.1,
)
trainer = TrainerClass(
model=model,
train_loader=DataLoader(train_ds, batch_size=512, collate_fn=collator),
val_loader=DataLoader(val_ds, batch_size=512, collate_fn=collator),
epochs=1000,
patience=5,
lr=1e-3,
criterion=nn.HuberLoss(),
)
trainer.fit()
ForecastingModel β embedding-based forecasting
from charm_toolkit import ForecastingModel, create_forecasting_datasets, collator, TrainerClass
from torch.utils.data import DataLoader
train_ds, val_ds, test_ds = create_forecasting_datasets(
raw_data,
descriptions=channel_names,
train_horizon=96,
test_horizon=96,
train_ratio=0.7,
val_ratio=0.15,
sequential=True,
scale=True,
)
model = ForecastingModel(
embedding_client=client,
horizon=96,
input_size=96,
head="linear",
hidden_dim=128,
mode="last", # "last", "avg", "none"
per_channel=True,
num_channels=len(channel_names),
)
trainer = TrainerClass(
model=model,
train_loader=DataLoader(train_ds, batch_size=512, collate_fn=collator),
val_loader=DataLoader(val_ds, batch_size=512, collate_fn=collator),
epochs=1000,
patience=10,
lr=1e-2,
)
trainer.fit()
ClassificationModel β time series classification
from charm_toolkit import ClassificationModel, create_classification_datasets, collator, TrainerClass
from torch.utils.data import DataLoader
import torch.nn as nn
train_ds, val_ds, test_ds = create_classification_datasets(
raw_data, # (N, T, C)
labels=labels, # list of N integer labels
descriptions=channel_names,
train_ratio=0.7,
val_ratio=0.15,
)
model = ClassificationModel(
embedding_client=client,
num_classes=num_classes,
hidden_dim=128,
pooling_over_t="mean",
pooling_over_channels="mean",
classifier_type="mlp",
)
trainer = TrainerClass(
model=model,
train_loader=DataLoader(train_ds, batch_size=32, collate_fn=collator),
val_loader=DataLoader(val_ds, batch_size=32, collate_fn=collator),
epochs=100,
patience=10,
lr=1e-3,
criterion=nn.CrossEntropyLoss(),
)
trainer.fit()
Precomputing embeddings (critical for training)
Toolkit models call the API every forward pass. For training with hundreds of windows per epoch, precompute embeddings once:
from charm_toolkit import precompute_dataset_embeddings, PrecomputedEmbeddingsDataset
# Compute once, save to disk as memmap
train_shape = precompute_dataset_embeddings(
client=client, dataset=train_ds,
output_path="./outputs/train_embeddings.pt", memory_batch_size=8192
)
val_shape = precompute_dataset_embeddings(
client=client, dataset=val_ds,
output_path="./outputs/val_embeddings.pt", memory_batch_size=8192
)
# Wrap datasets β model skips API calls when "embeds" key present
train_ds = PrecomputedEmbeddingsDataset(train_ds, "./outputs/train_embeddings.pt", train_shape)
val_ds = PrecomputedEmbeddingsDataset(val_ds, "./outputs/val_embeddings.pt", val_shape)
# Training now uses cached embeddings β orders of magnitude faster
train_loader = DataLoader(train_ds, batch_size=512, shuffle=True, collate_fn=collator)
Trainer API
from charm_toolkit import TrainerClass
trainer = TrainerClass(
model=model,
train_loader=train_loader,
val_loader=val_loader,
test_loader=test_loader, # optional
lr=1e-3,
weight_decay=1e-4,
epochs=1000,
patience=5,
min_delta=1e-4,
max_grad_norm=5.0,
criterion=None, # defaults to MSELoss
)
trainer.fit()
test_loss = trainer.evaluate(test_loader)
Dataset factory functions
All return (train_dataset, val_dataset, test_dataset):
| Function | Input shape | Key args |
|---|---|---|
create_reconstruction_datasets(raw_data, ...) |
(T, C) | window_size, stride, train_ratio, val_ratio |
create_forecasting_datasets(raw_data, ...) |
(T, C) | train_horizon, test_horizon, stride, train_ratio, val_ratio |
create_classification_datasets(raw_data, labels, ...) |
(N, T, C) | train_ratio, val_ratio |
Reconstruction and forecasting expect a single long time series (T, C) split temporally. Classification expects pre-windowed (N, T, C).
collator
All DataLoaders using toolkit datasets require collator as the collate_fn:
from charm_toolkit import collator
# or equivalently:
from charm_toolkit.Datasets import collator
Embeddings as features
CHARM embeddings work as drop-in feature vectors for any sklearn model:
import numpy as np
from sklearn.ensemble import IsolationForest
from sklearn.linear_model import LogisticRegression
from charm_toolkit.retrieval import cosine_similarity_matrix
response = client.embeddings.create(
descriptions=descriptions,
ts_array=windows_list,
return_tensors="np",
)
X = response.embeds # (N, D)
# Anomaly detection with isolation forest
clf = IsolationForest(contamination=0.05)
anomaly_labels = clf.fit_predict(X)
# Similarity search
sim = cosine_similarity_matrix(X, X)
# As features for any classifier
clf = LogisticRegression().fit(X_train, y_train)
Local Deployment
Deploy models locally from GitHub releases β no remote server needed:
with CharmClient(tag="experiment-2026-03-15_10-30-00") as client:
response = client.embeddings.create(...)
# Server shuts down automatically
When tag is provided:
- Checks for GPU availability (falls back to CPU)
- Clones repo at the specified tag (shallow clone)
- Downloads model weights from the GitHub release
- Launches the serving stack locally
- Polls health endpoint until ready
Files cached at ~/.charm/models/<tag>/ for fast subsequent runs.
CharmClient(
tag="experiment-tag", # required for local mode
repo_url="https://...", # default: c3-e/research
cache_dir="/path/to/cache", # default: ~/.charm/models
port=8080, # 0 = auto-select
)
Best practices
Input & preprocessing
- Send data as
(N, T, C)β N series, each T timesteps Γ C channels; all series in one request must share the same T and C. - Keep
T β€ 8192(the training window). The model can accept more, but quality degrades on lengths it wasn't trained on. T does not need to be a multiple of the patch size β inputs are padded to a patch boundary internally. - Do not pre-scale your data. CHARM normalizes internally (asinh z-score); applying StandardScaler/MinMaxScaler/log yourself hurts results.
Channel descriptions (a real quality lever)
- Use meaningful, consistent channel names (
"engine_temperature", not"col_0") β the model is channel-aware and descriptions materially affect embeddings. - Reuse the same names across requests so embeddings stay comparable (retrieval, clustering).
Batching & throughput
- Respect the per-request budget:
batch_size Γ C Γ T β€ 500,000(SDK-enforced client-side). - Use
async_createfor large N β it batches concurrently; the sync client is sequential and slow past ~100 series. Tunemax_B_per_request/concurrency_per_callinstead of one giant request.
Timeouts & retries
- Raise
timeoutfor forecasting β the SDK default is 15s, but the server allows forecasts up to ~220s. Usetimeout β 220+forprediction.create. - Keep the built-in retries (exponential backoff on 429/5xx) rather than hand-rolling.
Embeddings
aggregate=True(default) β(N, D)for retrieval / classification / clustering.aggregate=Falseβ(N, T_, C, D)only when you need per-patch/per-channel detail for a custom head.- L2-normalize before cosine similarity (embeddings are unit-normed by the encoder, but normalize again after any pooling you do).
- Discover
Dat runtime viaclient.model_info()β it's model-dependent; don't hardcode.
Forecasting
target_len > 0= forecast,< 0= backcast;0is invalid.- Keep
T + abs(target_len) β€ 8192(the training window) for best forecast/backcast quality. - Use
response.medianfor a point forecast, or the full quantile axis for intervals.Qis model-dependent (e.g. 21 or 99) β readdenormalized_predictions.shape[-1].
Classification
- For classification heads, don't pool over channels β keeping the per-channel embeddings flat (rather than averaging them) boosts accuracy, at a modest cost in head size/complexity. In the toolkit
ClassificationModel, setpooling_over_channels="flatten"(and passnum_channels=C, required for flatten) instead of the default"mean".
Dual-model awareness
- Treat
/predict(embeddings) and/forecastas separate models β they may have different patch sizes / embedding dims. Read the per-rolemodelsmap fromclient.model_info()instead of assuming they match.
Training on top of CHARM
- Precompute embeddings once (
precompute_dataset_embeddings+PrecomputedEmbeddingsDataset) β toolkit models otherwise call the API every forward pass.
Reliability & ops
- Catch specific errors (
InvalidRequestError,AuthenticationError,RateLimitError, or the baseCharmError) and back off on rate limits. - Use the context manager (
with CharmClient(...) as client:) so local deployments shut down cleanly. - Set credentials via env (
CHARM_API_KEY,CHARM_BASE_URL) rather than hardcoding.
Decision guide
When to use CHARM
- Multivariate time series (multiple channels measured over time)
- Each window has at least a few patches (patch size is 16, so ~48+ timesteps is a good floor)
- You want a strong starting point without feature engineering
When to use classical methods instead
- Tabular data without a time dimension β use LightGBM, XGBoost
- Very short series (< 10 timesteps)
- Single scalar features β still works but may not outperform ARIMA/ETS
Zero-shot vs fine-tuned
| Approach | When | Effort |
|---|---|---|
prediction.create(target_len=H) |
Quick forecast baseline, no labeled data | None β one API call |
| Embeddings + sklearn | Moderate data, combine with other features | Minutes |
| Embeddings + kNN (retrieval/AD) | Unlabeled anomaly detection or search | Minutes |
| Toolkit model (Reconstruction/Forecasting/Classification) | Have labeled data, want best performance | Train a small head (~minutes on CPU) |
GIFT-Eval results
Evaluated zero-shot on the full GIFT-Eval benchmark (97 dataset/frequency/term configurations) using the standard 11-metric protocol. Aggregate scores (geometric mean of per-config metrics normalized to the Seasonal Naive baseline; lower is better):
| Metric | Score (rel. Seasonal Naive) |
|---|---|
| MASE | 0.7582 |
| CRPS (mean weighted sum quantile loss) | 0.4776 |
Per-term (geometric mean, normalized to Seasonal Naive):
| Term | MASE | CRPS |
|---|---|---|
| short | 0.7463 | 0.5036 |
| medium | 0.7577 | 0.4452 |
| long | 0.7911 | 0.4460 |
Scores are the geometric mean of per-config metric / Seasonal Naive across all
97 GIFT-Eval configurations (lower is better; < 1.0 beats Seasonal Naive).
Full per-config results are in all_results.csv.
Evaluation protocol
- Benchmark: GIFT-Eval, 97 configs (short / medium / long terms).
- Metrics: MSE[mean], MSE[0.5], MAE[0.5], MASE[0.5], MAPE[0.5], sMAPE[0.5],
MSIS, RMSE[mean], NRMSE[mean], ND[0.5], mean_weighted_sum_quantile_loss
(computed with gluonts
evaluate_forecasts). - Context length: 8192; forecasts are full quantile distributions.
- Zero-shot: no GIFT-Eval train/test data is seen during pretraining
(
testdata_leakage = No).
Limitations
- Forecast quality varies by domain and horizon; very long horizons and highly non-stationary series remain challenging.
- Quantile calibration is learned and may drift on out-of-distribution scales.
Citation
@misc{charm,
title = {CHARM: A Zero-Shot Time-Series Foundation Model},
author = {C3 AI},
year = {2026},
url = {https://huggingface.co/c3aiia3c/CHARM}
}