Build stock research

Build our stock-research pipeline. Land the market API’s news and daily prices in Snowflake, clean them with dbt, and use Python to extract LLM signals alongside our numerical trend model. Serve the results to BI through BigQuery.

I’ll inspect the ingestion job and Google Docs research brief first. News sentiment and numerical forecasts will keep separate provenance, then meet in a point-in-time research mart.

The sample has 12 article revisions. I’ll retain the latest version available at the research cutoff—not future revisions. LLM confidence describes extraction certainty, not a stock-price probability. Should low-confidence labels be excluded?

Keep them. Flag anything below 0.80 for review, and don’t lose events if the model fails. Run it in dev before we publish.

The dev rehearsal is ready: 228 unique events, with 211 ready for reporting, 15 flagged for review, and 2 retained with analysis errors. All seven dbt checks passed. The plan and notebook are open on the right; nothing has been published to production. BigQuery will receive separate semantic-signal and trend-forecast views for the research dashboard.

news, prices = market_api.snapshot(as_of=cutoff)
warehouse.append_raw(news, "RAW.MARKET.NEWS")
warehouse.append_raw(prices, "RAW.MARKET.DAILY_BARS")
RAW.MARKET
dbt build --select int_news_by_ticker int_price_features
# Source URLs and as_of_at remain attached to every row.
Point-in-time inputs
Structured output
event
Product launch
sentiment
Positive
confidence
0.92
evidence
Source article retained
review_status
Ready
Structured signals
features = warehouse.table("INT_PRICE_FEATURES")
forecast = model.predict(features, horizon="5d")
forecast.save("RESEARCH.TREND_FORECASTS")
Trend forecasts
dbt build --select fct_stock_research
# Publishing requires approval; no production writes in this demo.
BigQuery → BI