TempatGunting

Feature Store Architecture for Production ML Pipelines

Feature engineering consumes 60-80% of the effort in production machine learning projects. Teams compute features in notebooks for training, rewrite them in a different language for serving, discover inconsistencies between the two implementations months later, and debug models that mysteriously perform worse in production than in evaluation. Feature stores exist to eliminate these failure modes by providing a single source of truth for feature definitions, computation, and serving.

A feature store is not a database — it is an abstraction layer over multiple storage systems that serves features at different latencies for different use cases. The offline store provides historical feature values for training and batch prediction. The online store provides the latest feature values for real-time prediction. The feature store ensures that features are computed identically for both, eliminating the training-serving skew that silently degrades model performance.

Dual-Store Architecture

The core architecture splits feature storage into two tiers optimized for their respective access patterns. The offline store is optimized for analytical queries — scanning large time ranges across many entities for training dataset construction. The online store is optimized for point lookups — retrieving the latest feature vector for a single entity during real-time inference.

# Feast feature store definition
from feast import FeatureStore, Entity, FeatureView, Field
from feast.types import Float64, Int64, String
from feast.infra.offline_stores.bigquery_source import BigQuerySource

user = Entity(name="user", join_keys=["user_id"])

user_activity = FeatureView(
    name="user_activity_features",
    entities=[user],
    schema=[
        Field(name="total_purchases_30d", dtype=Int64),
        Field(name="avg_order_value_30d", dtype=Float64),
        Field(name="days_since_last_purchase", dtype=Int64),
        Field(name="favorite_category", dtype=String),
        Field(name="session_count_7d", dtype=Int64),
    ],
    source=BigQuerySource(
        table="ml_features.user_activity_daily",
        timestamp_field="feature_timestamp",
        created_timestamp_column="created_at",
    ),
    online=True,
    ttl=timedelta(days=1),
)

The online=True flag tells Feast to materialize this feature view to the online store. During materialization, Feast reads the latest feature values from the offline store (BigQuery) and writes them to the online store (Redis or DynamoDB). The ttl specifies how long feature values remain valid — stale values older than the TTL are not returned during online serving.

Point-in-Time Joins

Point-in-time correctness is the most important property a feature store provides. When constructing a training dataset, features must be joined to training events using the timestamp of each event, not the current time. This prevents future data leakage — a subtle bug that inflates evaluation metrics while degrading production performance.

# Point-in-time correct training dataset
from feast import FeatureStore

store = FeatureStore(repo_path="./feature_repo")

# Training events with timestamps
entity_df = pd.DataFrame({
    "user_id": [101, 102, 103, 101, 102],
    "event_timestamp": pd.to_datetime([
        "2026-09-01", "2026-09-05", "2026-09-10",
        "2026-09-15", "2026-09-20"
    ]),
    "label": [1, 0, 1, 0, 1]
})

# Feast joins features as-of each event timestamp
training_df = store.get_historical_features(
    entity_df=entity_df,
    features=[
        "user_activity_features:total_purchases_30d",
        "user_activity_features:avg_order_value_30d",
        "user_activity_features:days_since_last_purchase",
    ]
).to_df()

The get_historical_features call joins each event row with the feature values that were available at that event's timestamp. For user 101's September 1st event, it retrieves features computed before September 1st. For the same user's September 15th event, it retrieves features computed before September 15th. This as-of join ensures the training data accurately reflects what the model would see in production.

Feature Computation Patterns

Batch Features

Batch features are computed on a schedule — hourly, daily, or triggered by upstream data arrival. They aggregate historical data into feature values that change slowly. Examples include 30-day purchase counts, average session duration, and user segment classifications. Batch computation runs in the data warehouse using SQL or Spark and writes results to the offline store.

# Batch feature computation in SQL (runs daily)
CREATE OR REPLACE TABLE ml_features.user_activity_daily AS
SELECT
    user_id,
    COUNT(*) FILTER (WHERE order_date >= CURRENT_DATE - 30) AS total_purchases_30d,
    AVG(order_total) FILTER (WHERE order_date >= CURRENT_DATE - 30) AS avg_order_value_30d,
    CURRENT_DATE - MAX(order_date) AS days_since_last_purchase,
    MODE() WITHIN GROUP (ORDER BY category) AS favorite_category,
    COUNT(DISTINCT session_id) FILTER (
        WHERE session_start >= CURRENT_DATE - 7
    ) AS session_count_7d,
    CURRENT_TIMESTAMP AS feature_timestamp,
    CURRENT_TIMESTAMP AS created_at
FROM orders
JOIN sessions USING (user_id)
GROUP BY user_id;

Streaming Features

Streaming features update in near real-time from event streams. They capture recent behavior that batch features miss — clicks in the last 5 minutes, items currently in the cart, and real-time session engagement scores. Streaming computation uses Flink, Spark Structured Streaming, or Kafka Streams and writes directly to the online store. The design considerations overlap with those in Apache Flink ML pipelines.

# Streaming feature computation with Flink SQL
CREATE TABLE user_realtime_features (
    user_id STRING,
    clicks_5min INT,
    cart_value DOUBLE,
    session_duration_sec INT,
    feature_timestamp TIMESTAMP(3),
    WATERMARK FOR feature_timestamp AS feature_timestamp - INTERVAL '10' SECOND
) WITH (
    'connector' = 'upsert-kafka',
    'topic' = 'user-features',
    'key.format' = 'json',
    'value.format' = 'json'
);

INSERT INTO user_realtime_features
SELECT
    user_id,
    COUNT(*) OVER w AS clicks_5min,
    SUM(item_price) OVER w AS cart_value,
    TIMESTAMPDIFF(SECOND, MIN(event_time) OVER w, MAX(event_time) OVER w) AS session_duration_sec,
    MAX(event_time) OVER w AS feature_timestamp
FROM clickstream
WINDOW w AS (
    PARTITION BY user_id
    ORDER BY event_time
    RANGE BETWEEN INTERVAL '5' MINUTE PRECEDING AND CURRENT ROW
);

Online Serving

Online serving retrieves feature vectors for real-time model inference. The critical requirement is latency — feature retrieval must complete within the model serving SLA, typically under 10 milliseconds at p99. This constrains the online store to key-value stores with single-digit millisecond read latency.

# Online feature retrieval during inference
from feast import FeatureStore

store = FeatureStore(repo_path="./feature_repo")

# Retrieve features for a single user at serving time
features = store.get_online_features(
    features=[
        "user_activity_features:total_purchases_30d",
        "user_activity_features:avg_order_value_30d",
        "user_activity_features:days_since_last_purchase",
        "user_realtime_features:clicks_5min",
        "user_realtime_features:cart_value",
    ],
    entity_rows=[{"user_id": "101"}]
).to_dict()

# Pass to model for inference
prediction = model.predict(features)

Feature retrieval latency depends on the online store backend and the number of features requested. Redis returns single feature vectors in under 1ms. Retrieving features from multiple feature views requires multiple store reads that can be parallelized. For high-throughput serving, batch multiple entity lookups into a single request using get_online_features with multiple entity rows.

Feature Store Comparison

CapabilityFeastTectonHopsworks
LicenseOpen source (Apache 2.0)CommercialOpen source + enterprise
Offline storeBigQuery, Snowflake, Redshift, fileDatabricks, Snowflake, S3Hudi on S3/HDFS
Online storeRedis, DynamoDB, SQLiteDynamoDB, RedisRonDB (MySQL NDB)
Streaming featuresVia push APINative Spark/FlinkNative Spark/Flink
Feature monitoringCommunity pluginsBuilt-inBuilt-in
Best forSmall-medium teamsEnterprise ML platformsData-intensive ML

Feast is the right starting point for most teams. It provides the core feature store primitives — feature definitions, point-in-time joins, online materialization, and a feature registry — without the operational complexity of a managed platform. Graduate to Tecton or Hopsworks when you need native streaming feature computation, built-in monitoring, or enterprise features like access control and feature lineage.

Avoiding Common Mistakes

Building features in notebooks and translating them to SQL for production introduces training-serving skew through subtle differences in null handling, type casting, and aggregation boundaries. Define features in the feature store from the start, even if the initial implementation is simple. The feature store's compute engine ensures identical behavior across training and serving.

Materializing all features to the online store wastes resources and increases materialization latency. Only materialize features that are actually used in real-time inference. Features used exclusively for batch scoring or training do not need online materialization. Audit online store usage quarterly and remove feature views that are materialized but never queried, applying the same analysis mindset used in vector database performance tuning.

Ignoring feature freshness requirements leads to models making predictions on stale data. Define a freshness SLA for each feature view — how old a feature value can be before it degrades model performance. Monitor actual freshness against the SLA and alert when materialization lag exceeds the threshold. A recommendation model using 24-hour-old click features performs very differently from one using 5-minute-old features.