PLANNEDMachine Learning · Statistical Research · Futures

Machine Learning and Textual Information for Trend Research

A planned research stream on whether textual information improves the timing, filtering, or interpretation of systematic trend signals.

Started Not startedSample [SAMPLE PERIOD TO CONFIRM]Python · pandas · scikit-learn · statsmodels
02

Research question

Can textual information improve the timing, filtering, or interpretation of systematic trend signals beyond what price, volatility, and liquidity variables already provide?

03

Why this matters

Trend signals treat all price moves alike. If textual information can distinguish moves accompanied by genuine new information from moves that are mostly noise, signal filtering could improve without changing the underlying trend logic. The risk is that flexible models and small effective sample sizes make overfitting almost effortless.

04

Hypotheses

  1. 01H1: Textual features can distinguish information-driven trends from noisy price moves at a horizon relevant to trend signals.
  2. 02H2: NLP variables improve trend-signal filtering without degrading out-of-sample performance relative to a price-only baseline.
  3. 03H3: Textual features are incremental to price, volatility, and liquidity variables rather than proxies for them.
05

Methodology

Preregistration area

Hypotheses, candidate data, features, evaluation protocol, and stopping rules are written down before estimation begins, so that later results can be read against the original plan.

  • Fixed baseline: a price-only trend specification defined in advance.
  • Fixed evaluation: walk-forward windows declared before any model is fit.
  • Fixed feature list: additions after the fact are reported as exploratory, not confirmatory.

Overfitting controls

  • Strict separation of model selection and evaluation windows
  • Limits on the number of specifications tested, with correction for those tested
  • Preference for low-capacity models before flexible ones
  • Ablation tests to show incremental value over price-based variables
  • Reporting of every specification attempted, not only the surviving one

Proposed data

[TEXT DATA SOURCE TO CONFIRM] Licensing and point-in-time availability must be verified before any test, since text archives are frequently revised.

Preregistered hypotheses and specificationPrice-only baseline as the comparisonWalk-forward evaluationNested cross-validation for hyperparametersMultiple-testing correctionFeature-ablation testing
06

Data

Status

No data has been acquired. Any source used will be documented with coverage, point-in-time availability, and licensing constraints.

07

Key results

No results published yet

No results exist. This project is planned and no tests have been completed. Nothing on this page should be read as evidence.

Planned exhibits

  • Preregistered hypothesis table
  • Baseline versus augmented walk-forward comparison
  • Feature-ablation results
  • Specification-count and correction summary

Every published chart will carry a descriptive title, axis labels, units, legend where needed, its sample period, a gross or net label, a short written interpretation, and accessible colors with tooltips.

08

Interpretation

Status

Interpretation cannot precede estimation. This section will be written only after preregistered tests are run.

09

Robustness checks

  • Walk-forward window sensitivity
  • Alternative text representations
  • Ablation against price-only baseline
  • Multiple-testing corrections
  • Sample-period stability
10

Limitations

  • Text archives are often revised, creating look-ahead risk that is hard to detect.
  • Effective sample size at monthly horizons is small relative to model flexibility.
  • Language coverage differs sharply across markets and asset classes.
  • Positive findings in this area are frequently unreplicable.
11

Conclusion

No conclusion. The project is at the design stage and is published here to fix its hypotheses in advance.

12

Reproducibility

Plan

The preregistration document, once finalized, will be versioned in the repository with a timestamp before any estimation code is run.

13

Downloads and links

Read Full ReportpendingView CodependingOpen NotebookpendingDownload FigurespendingMethodology AppendixpendingData Dictionarypending
14

Citation

This is working research, not a peer-reviewed publication. If you refer to it, please cite it as work in progress and note the status shown above.

BibTeX

@misc{bang_ml_text_trend_research,
  author = {Bang, Pratik},
  title  = {Machine Learning and Textual Information for Trend Research},
  year   = {2026},
  note   = {Status: Planned. Working research, subject to revision.},
  url    = {[PROJECT URL TO ADD]}
}

Plain text

Bang, P. (2026). Machine Learning and Textual Information for Trend Research. Working research (Planned). [PROJECT URL TO ADD]