NLP-on-10Ks-from-EDGAR-DB
Unverified ML strategy on Multi by lucaskienast. BotFinder score 18 out of 100.
This is a sentiment trading strategy, written in Python, and applying NLP on 10-K's from the SEC EDGAR database.
Source: github
BotFinder analysis pending.
NLP-on-10Ks-from-EDGAR-DB
Algotrading NLP on 10-K's In this project, NLP Analysis was carried out on 10-k financial statements to generate an alpha factor. For the dataset, the end of day from Quotemedia and Loughran-McDonald sentiment word lists were used. Installation Use git clone to get a copy of this repository. Method - define list of public companies and get their CIK's - use BeautifulSoup to get XML files for each CIK - download complete submission text file for all CIK's with secedgardownloader - use re to get text content between tags where is '10-k' - use BeautifulSoup to remove tags and convert to lower case - lemmatize verbs using nltk.stem.WordNetLemmatizer - remove stopwords using nltk.corpus.stopwords - get Loughran-McDonald sentiment word list and lemmatize it - generate bag of words that counts number of sentiment words in each 10-K via sklearn.featureextraction.text.CountVectorizer - compute Term Frequency Inverse Document Frequency (TFIDF) via sklearn.featureextraction.text.TfidfVectorizer - compute cosine similarity to evaluate change in TFIDF over time using sklearn.metrics.pairwise.cosi
⚠ No verified equity curve — no track-record source connected.
Drawdown profile
Data unavailable — contact the owner.
Verification ledger
How the score has moved
Recalculated at each data collection. Transparency means showing the bad weeks too.
No score history is stored yet — only the current score is shown.
Reviews & comments
No reviews collected from the source yet.
⚠ No live verification account connected — ask for proof before buying.
Alerts on changes: coming soon
Prop-firm compatibility not provided.
Open-source maintainer on GitHub.
Data-completeness & trust index (not a profitability rating)