Data is an operational input, not just a chart
A trading bot can only evaluate the information it receives. Depending on its task, it may consume price bars, individual trades, volume, quotes, order-book changes, market status, reference data, account balances, or signals from another process. A data feed is therefore part of the operating system around the bot, not simply a display of past prices.
The relevant question is not “which data is best?” in the abstract. It is whether the chosen observations are appropriate for the bot's purpose, available when expected, consistently represented, and handled safely when something goes wrong. An alerting tool and a bot reacting to rapidly changing quotes can have very different freshness requirements.
This article focuses on inputs and their operational quality. The wider Trading Bots hub covers the cluster, while Trading Bot Architecture shows where data validation fits into the rest of a system.
Common forms of market data
Price data can be delivered as periodic bars or as individual events. OHLCV bars summarize the open, high, low, close, and volume over a defined interval. They are compact and convenient for many monitoring tasks, but a bar hides the order and timing of events inside that interval. A bar's timestamp convention—start time, end time, or publication time—also affects how a bot should interpret it.
Trade or tick data records individual reported transactions or price updates, depending on the provider's definition. It can show event-level changes but creates more records and may require careful handling of corrections, duplicates, and sequencing. Quote data describes available bid and ask prices and quantities; a spread can be calculated from the bid and ask, but the displayed quote may change before an order reaches the venue.
Order-book data represents resting interest at one or more price levels. It can be a snapshot or a stream of changes that must be applied in order to maintain a current view. Data coverage and depth vary by source and venue. A partial book should not be mistaken for a complete picture of market interest.
A bot may also need instrument metadata, trading calendars, status flags, currency or unit conventions, and account state. Those fields can affect whether a price is valid, a market is open, or a requested quantity can be represented. The required set follows from the task and execution venue rather than from a universal checklist.
Timestamps, ordering, and freshness
A timestamp can refer to when an event occurred, when a provider received it, or when the bot consumed it. These are not interchangeable. If source and local clocks differ, apparent ordering may be misleading. Systems should document which timestamp they use, the relevant timezone, and how clock differences are handled.
Freshness is the age of an observation when the bot uses it. A feed can be connected yet stale if updates stop or are delayed. Whether an observation is too old depends on the intended workflow, so freshness limits should be tied to the task and monitored explicitly. There is no single time threshold appropriate for every market and strategy.
Ordering also matters. A late event can arrive after a newer event; a reconnect may replay prior messages; two sources may report related changes at different times. The receiving process needs a defined policy for duplicates, out-of-order information, and gaps. If the system cannot establish a coherent current state, passing data onward as though it were complete can create silent errors.
Missing, duplicated, and inconsistent observations
Missing data can arise from a connection interruption, provider gap, inactive instrument, market halt, or a legitimate absence of activity. These cases have different meanings. Replacing a missing value with the last known price may be acceptable for one display but misleading for another decision; the policy should be explicit and appropriate to the input.
Duplicate observations can appear during retries, reconnects, or provider replay. Processing the same event twice may cause repeated calculations or, in a poorly designed workflow, duplicate downstream actions. Event identifiers, sequence information, or other checks can help identify repeats, but the correct method depends on the source.
Inconsistencies can include different symbol formats, precision, currency units, calendar rules, or bar boundaries between feeds. Normalizing data means mapping these conventions into an understood internal representation; it does not make sources equivalent. Preserve enough provenance to know where each value came from and which transformations were applied.
When a quality check fails, the bot should have a deliberate response. It might flag the record, hold the latest valid state with a freshness warning, pause a downstream action, or request human review. The right response depends on the consequences of continuing. Silently filling gaps or ignoring validation errors can make a system appear healthy when its decision inputs are not.
Market coverage and source differences
A data source describes only the venues, instruments, periods, and fields it covers. Prices, volume, and order-book depth can differ across venues, and some feeds aggregate information while others report venue-specific events. A bot using one source should not assume it sees every relevant market or that its view is the same as the execution venue's view.
Historical and live services can also differ in format, filtering, correction policies, and timestamps. A strategy may be developed using one dataset and deployed against another stream. Before relying on a handoff, compare schemas and conventions and identify which differences matter to downstream logic.
Coverage questions include whether an instrument is active, whether a feed includes all expected trading sessions, how corporate or instrument events are represented, and what happens during maintenance. These details are especially important when a bot monitors a portfolio or a set of markets with different calendars.
Historical data and live data serve different roles
Historical data is used to inspect earlier observations, develop logic, or evaluate behavior over a past period. Live data arrives during operation and can have latency, interruptions, and venue-specific timing that a stored file may not reproduce. A bot may use both, but the transition from historical research to a live input path should be deliberate.
A simulation based on historical bars may not represent the sequence of quotes or order-book changes that would have been available at each moment. Likewise, a live feed may include delays, corrections, or gaps not apparent in a clean research dataset. The data source and assumptions should be recorded so that test results are not confused with the behavior of the deployed feed.
This article is about data used by the bot, not the full evaluation method. For model-specific construction of inputs from raw observations, see feature engineering in machine learning. That topic concerns transforming information into model features; operational normalization here concerns the integrity, timing, and representation of the input stream.
Handoff from feed to strategy or model
A bot's data layer often transforms provider messages into a structure that downstream logic can consume. It might parse fields, align timestamps, calculate simple summaries, attach instrument identifiers, and publish an event or current-state snapshot. The handoff should define what each field means, its units, its time reference, and how missing values are represented. When the feed is supplied through a venue interface, the trading bot API guide explains the request and streaming patterns that can carry it.
The strategy or model should be able to tell whether its required inputs are present and current. If the data contract changes, an input may still parse while carrying a different meaning. Versioned schemas, validation checks, and tests using known examples can help detect that kind of mismatch before it affects live behavior.
For a model, the expected input features may be more processed than the raw feed. Those transformations belong to the model pipeline and need their own consistency checks. Keeping feed validation separate from model feature construction makes it easier to locate whether a problem originates in collection, normalization, or later interpretation.
Monitor data quality as part of bot operation
Monitoring should make data failures visible. Useful checks may include last-update time, missing-field counts, sequence gaps, duplicate rates, unexpected values, and whether expected instruments remain covered. The measures should match the source and intended use; an alert should point to the feed or field that needs investigation rather than merely report that the bot is online.
Consider a hypothetical bot that watches a market-price feed. If updates stop but the connection remains open, a simple “connected” indicator might stay green. A freshness check can identify that the last observation is older than the workflow allows. The bot could then stop preparing new actions, record the interruption, and alert its operator until valid updates resume. This example illustrates an operational safeguard, not a recommendation about a particular market or threshold.
Logs should record enough information to compare source events with downstream decisions, including relevant timestamps and validation outcomes. When correcting a feed problem, preserve the distinction between what the bot observed and what was later repaired or backfilled. Otherwise, a review may incorrectly attribute a decision to information that was not available at that time. These controls are part of the broader trading technology stack.
A practical data review
Before connecting a feed to a bot, document the source, supported instruments, fields, timestamps, expected update behavior, and known coverage limits. Test normal observations as well as missing values, repeated messages, delayed updates, reconnects, and unexpected formats. Confirm that the system's response to each case is observable.
Then verify the handoff to the next component. Check that a strategy or model receives the expected values with the intended units and timing, and that invalid or stale data cannot quietly pass as current. Where multiple feeds are used, define which source takes priority and how disagreement is handled.
Market data quality is not a guarantee of a good decision; it is one condition for the bot to operate as designed. Even clean, timely data can be incomplete for a purpose or fail to capture future conditions. Treat input monitoring as part of the system lifecycle, alongside order and account-state checks described in how AI trading bots work. The broader bot testing guide includes checks for feed interruptions and stale input; bot risk management and operational monitoring cover how those problems are controlled and observed after deployment.