news, prices = market_api.snapshot(as_of=cutoff)
warehouse.append_raw(news, "RAW.MARKET.NEWS")
warehouse.append_raw(prices, "RAW.MARKET.DAILY_BARS")
Build stock research
I’ll inspect the ingestion job and Google Docs research brief first. News sentiment and numerical forecasts will keep separate provenance, then meet in a point-in-time research mart.
1news, bars = market_api.snapshot(as_of=research_cutoff)2warehouse.append_raw(news, "RAW.MARKET.NEWS")3warehouse.append_raw(bars, "RAW.MARKET.DAILY_BARS")| 1 | event_id | string |
|---|---|---|
| 2 | ticker | string |
| 3 | source_url | string |
| 4 | published_at | string |
| 5 | updated_at | string |
| 6 | body | string |
240 rows · 228 event IDs. Twelve revisions share an existing ID; keep the latest revision available at the research cutoff before inference.
Equity Research Playbook
Use the approved security universe and point-in-time inputs. Keep source URLs and review anything below 0.80 confidence. The numerical forecast is separate from the LLM’s semantic analysis.
Read prepared warehouse tables in Python for analysis, then write the results back for downstream dbt models.
The sample has 12 article revisions. I’ll retain the latest version available at the research cutoff—not future revisions. LLM confidence describes extraction certainty, not a stock-price probability. Should low-confidence labels be excluded?
+1Three files ready for review
enrich_news.py extracts semantic signals from dbt-prepared news. fct_stock_research.sql joins those signals to the existing numerical forecasts. The YAML tests the research grain and score ranges before publishing separate views to BigQuery.
Rehearsal complete
240 source rows → 228 unique events → 211 ready, 15 needing review, and 2 endpoint errors. Seven dbt checks passed. Errors remain in the mart with null confidence. The numerical model keeps its existing walk-forward validation; this run makes no prediction-quality claim.
The dev rehearsal is ready: 228 unique events, with 211 ready for reporting, 15 flagged for review, and 2 retained with analysis errors. All seven dbt checks passed. The plan and notebook are open on the right; nothing has been published to production. BigQuery will receive separate semantic-signal and trend-forecast views for the research dashboard.
dbt build --select int_news_by_ticker int_price_features
# Source URLs and as_of_at remain attached to every row.
- event
- Product launch
- sentiment
- Positive
- confidence
- 0.92
- evidence
- Source article retained
- review_status
- Ready
features = warehouse.table("INT_PRICE_FEATURES")
forecast = model.predict(features, horizon="5d")
forecast.save("RESEARCH.TREND_FORECASTS")
dbt build --select fct_stock_research
# Publishing requires approval; no production writes in this demo.
Ingest market data
news, prices = market_api.snapshot(as_of=cutoff)
warehouse.append_raw(news, "RAW.MARKET.NEWS")
warehouse.append_raw(prices, "RAW.MARKET.DAILY_BARS")
Prepare research inputs
dbt build --select int_news_by_ticker int_price_features
# Source URLs and as_of_at remain attached to every row.
Extract news signals
- event
- Product launch
- sentiment
- Positive
- confidence
- 0.92
- evidence
- Source article retained
- review_status
- Ready
Model price trends
features = warehouse.table("INT_PRICE_FEATURES")
forecast = model.predict(features, horizon="5d")
forecast.save("RESEARCH.TREND_FORECASTS")
Build research mart
dbt build --select fct_stock_research
# Publishing requires approval; no production writes in this demo.
1from research.runtime import warehouse, llm, save_analysis2from research.policy import review_status3 4# Ingestion lands licensed API responses in Snowflake RAW.MARKET.5# dbt prepares point-in-time news and numerical price features.6news = warehouse.table("ANALYTICS.INTERMEDIATE.INT_NEWS_BY_TICKER")7 8for batch in news.to_pandas_batches():9 for row in batch.itertuples():10 result = llm.extract(11 articles=row.ARTICLES,12 fields={"topic": str, "sentiment_score": float,13 "rationale": str, "confidence": float},14 prompt_version="equity-v3",15 )16 # The helper validates ranges and retains failed model calls.17 save_analysis(18 table="ANALYTICS.RESEARCH.NEWS_SIGNALS",19 security_id=row.SECURITY_ID,20 as_of_at=row.AS_OF_AT,21 result=result,22 review_status=review_status(result, threshold=0.80),23 preserve_sources=True,24 )25 26# predict_trends.py separately uses lagged returns and volume.27# Its walk-forward forecast is not the LLM's confidence score.1from research.runtime import warehouse, llm, save_analysis2from research.policy import review_status3 4# Ingestion lands licensed API responses in Snowflake RAW.MARKET.5# dbt prepares point-in-time news and numerical price features.6news = warehouse.table("ANALYTICS.INTERMEDIATE.INT_NEWS_BY_TICKER")7 8for batch in news.to_pandas_batches():9 for row in batch.itertuples():10 result = llm.extract(11 articles=row.ARTICLES,12 fields={"topic": str, "sentiment_score": float,13 "rationale": str, "confidence": float},14 prompt_version="equity-v3",15 )16 # The helper validates ranges and retains failed model calls.17 save_analysis(18 table="ANALYTICS.RESEARCH.NEWS_SIGNALS",19 security_id=row.SECURITY_ID,20 as_of_at=row.AS_OF_AT,21 result=result,22 review_status=review_status(result, threshold=0.80),23 preserve_sources=True,24 )25 26# predict_trends.py separately uses lagged returns and volume.27# Its walk-forward forecast is not the LLM's confidence score.1{{ config(materialized='table') }}2 3with signals as (4 select * from {{ source('research', 'news_signals') }}5), forecasts as (6 select * from {{ source('research', 'trend_forecasts') }}7)8 9select10 coalesce(s.security_id, f.security_id) as security_id,11 coalesce(s.as_of_at, f.as_of_at) as as_of_at,12 s.sentiment_score,13 s.confidence,14 s.review_status,15 s.source_urls,16 s.prompt_version,17 f.predicted_return,18 f.forecast_horizon,19 f.training_cutoff,20 f.model_version21from signals s22full outer join forecasts f23 on s.security_id = f.security_id24 and s.as_of_at = f.as_of_at25 26-- Publish separate semantic and forecast views to BigQuery.27-- Keep review states and numerical model provenance distinct.1version: 22 3sources:4 - name: research5 database: ANALYTICS6 schema: RESEARCH7 tables:8 - name: news_signals9 - name: trend_forecasts10 11models:12 - name: fct_stock_research13 description: >14 Point-in-time equity research, served to BI through BigQuery.15 LLM confidence is not a probability of a stock-price move.16 data_tests:17 - dbt_utils.unique_combination_of_columns:18 arguments:19 combination_of_columns: [security_id, as_of_at]20 columns:21 - name: security_id22 data_tests: [not_null]23 - name: as_of_at24 data_tests: [not_null]25 - name: confidence26 data_tests:27 - dbt_utils.accepted_range:28 arguments:29 min_value: 030 max_value: 131 - name: sentiment_score32 data_tests:33 - dbt_utils.accepted_range:34 arguments:35 min_value: -136 max_value: 1