diff --git a/context/CONVENTIONS.md b/context/CONVENTIONS.md index 7f95e7e..cedb69c 100644 --- a/context/CONVENTIONS.md +++ b/context/CONVENTIONS.md @@ -10,6 +10,8 @@ Terse imperative code rules. No rationale here — rationale lives in - Check formatting: `uv run ruff format --check .` - Lint Markdown: `npx markdownlint-cli2 "**/*.md"` - Sync project dependencies: `uv sync` +- Do not add automated test suites, fixtures, or eval workspaces; validate with lint and focused + smoke commands. ## Python diff --git a/context/DECISIONS.md b/context/DECISIONS.md index ac94d94..63da46e 100644 --- a/context/DECISIONS.md +++ b/context/DECISIONS.md @@ -17,6 +17,21 @@ Status: active | superseded by --- +## 2026-08-31 — Distill routine 13F queries while preserving SEC-level provenance + +Context: SEC 13F filings are authoritative but awkward for reverse stock-holder lookup and +routine manager history, while 13f.info already resolves managers, CUSIPs, filing portfolios, +and position histories. Exposing provider URLs or low-level CUSIP flags would invite agents to +leave the tool surface and repeat work. +Decision: Use 13f.info behind the SEC skill's routine stock-, manager-, and manager-position +queries. Keep the agent-facing interface job-based, abstract the intermediary from generated +reports, and expose the underlying SEC period, CIK, and accession. Use raw EDGAR for fields the +distilled route omits, provider failure, discrepancies, or requested verification. +Tradeoff: Routine queries depend on an unofficial distilled backend and its HTML/JSON shape, but +agents get a smaller, more capable interface while retaining a direct path to the regulatory +record. +Status: active + ## 2026-08-12 — Replace Pi's coding identity without duplicating tool or skill guidance Context: A profile-level `SYSTEM.md` replaces Pi's default prompt, which removes the diff --git a/context/MAP.md b/context/MAP.md index 5c89362..b329548 100644 --- a/context/MAP.md +++ b/context/MAP.md @@ -51,7 +51,8 @@ sessions, and unrelated settings remain profile-local and untouched. `scan_conferences.py`, `search_themes.py`; universe config in `screens.json`; shared bootstrap in `scripts/_common.py`. - `skills/sec-edgar-skill/` — `scripts/fetch_*.py`, `parse_financials.py`, `orient.py`, - `list_headings.py`; guides in `references/`; shared bootstrap in `scripts/_common.py`. + `list_headings.py`; filing, financial, ownership, proxy/governance, and holdings guides in + `references/`; shared bootstrap in `scripts/_common.py`. - `skills/market-scout/` — `scripts/fetch_market_data.py`, `fetch_transcripts.py`, and shared `scripts/_common.py`. - `skills/bottom-up-analyst/` — valuation `scripts/dcf.py`, `epv.py`; archetypes and guides in @@ -71,8 +72,10 @@ flowchart TD - `bottom-up-analyst` is the brain and conductor: it decides what to pull, reasons over it, and writes the memo. The data skills never decide what matters. - The two filing/market data skills know nothing of each other and are swappable. -- `signal-sweep` and `sec-edgar-skill` both read `EDGAR_IDENTITY` and share on-disk cache - contracts defined in their `_common.py` modules. +- `signal-sweep` and SEC-facing `sec-edgar-skill` commands read `EDGAR_IDENTITY` and use + on-disk cache contracts defined in their `_common.py` modules. The SEC skill's routine 13F + convenience queries use 13f.info and expose underlying SEC periods and identifiers; raw + EDGAR remains the deep-field and verification route. ## Production order diff --git a/skills/README.md b/skills/README.md index 78bb7c9..12396a3 100644 --- a/skills/README.md +++ b/skills/README.md @@ -92,6 +92,7 @@ cloc \ skills/sec-edgar-skill/references/guide_filings.md \ skills/sec-edgar-skill/references/guide_financials.md \ skills/sec-edgar-skill/references/guide_ownership.md \ + skills/sec-edgar-skill/references/guide_proxy.md \ skills/sec-edgar-skill/references/guide_holdings.md \ skills/sec-edgar-skill/scripts/orient.py \ skills/sec-edgar-skill/scripts/fetch_filing.py \ diff --git a/skills/sec-edgar-skill/README.md b/skills/sec-edgar-skill/README.md index e537254..a46de2e 100644 --- a/skills/sec-edgar-skill/README.md +++ b/skills/sec-edgar-skill/README.md @@ -1,65 +1,70 @@ # SEC EDGAR Research Skill -A tools skill that teaches an AI coding agent how to retrieve and extract data from -**SEC EDGAR** filings for US-listed companies (domestic issuers and foreign private -issuers), efficiently and within a token budget. - -It is the **data/tools layer** of a research stack: it fetches and extracts; it does not -decide what matters. Pair it with an analytical-framework skill (which supplies the -judgment and the output shape) and, optionally, a presentation/consumer skill. Keeping -this layer unopinionated lets any framework compose on top of it. - -## What's included - -- **`SKILL.md`** — the entry point: setup, the token-efficient retrieval method, the - cache contract, and routing to the guides and scripts. -- **`references/`** — modular, lazily-loaded guides, one per data domain: - - `guide_core.md` — company lookup, filing discovery, `.to_context()`, `.docs`. - - `guide_filings.md` — filing text by SEC item code (10-K/10-Q/8-K/20-F) or heading discovery (DEF 14A/6-K); attachments (6-K Exhibit 99.1). - - `guide_financials.md` — XBRL statements and facts (US-GAAP & IFRS). - - `guide_ownership.md` — insider transactions (3/4/5) and executive compensation (DEF 14A; 20-F Item 6). - - `guide_holdings.md` — 13F institutional holdings and 13D/13G blockholders. -- **`scripts/`** — thin, self-documenting wrappers around `edgartools` (shared setup - lives in `_common.py`): - - `orient.py` — company summary + filing-mix survey + recent filings (run first). - - `fetch_filing.py`, `fetch_filings.py` — filings (and sections/attachments) to Markdown. - - `parse_financials.py` — XBRL statements to CSV. - - `list_headings.py` — heading→line map for a cached filing. - - `fetch_insider_trades.py` — insider transactions (Form 4 buys/sells). - - `fetch_13f_holders.py` — institutional 13F holders (via 13f.info). - - `test_setup.py` — environment diagnostics. - -## Setup - -1. **Install this skill's dependencies** from this directory: `uv sync` -2. **Set `EDGAR_IDENTITY`** — see [profile setup](../../README.md#one-time-runtime-setup). -3. **Verify:** `python scripts/test_setup.py --live` - -## Add the skill to your agent +An agent skill for retrieving, verifying, and inspecting SEC filings and filing-derived +financial, ownership, compensation, and governance data without loading whole filings into +context. + +## Included resources + +- `SKILL.md` — agent workflow, source hierarchy, cache contract, and resource routing. +- `references/guide_core.md` — company and filing discovery. +- `references/guide_filings.md` — sections, free-form filings, and exhibits. +- `references/guide_financials.md` — XBRL statements, facts, and reporting periods. +- `references/guide_ownership.md` — Forms 3/4/5 and insider transactions. +- `references/guide_proxy.md` — beneficial ownership, governance, compensation, and voting. +- `references/guide_holdings.md` — 13F and 13D/13G holdings. +- `scripts/` — command-line retrieval and extraction helpers. + +## Install this skill + +From the SecStack repository: + +```bash +npx skills add eggmasonvalue/secstack --skill sec-edgar-skill +``` + +SecStack's profile bootstrap installs the Python dependencies automatically. For a standalone +installation, run from this directory: + +```bash +uv sync +``` + +Then invoke the scripts with that environment, or configure the harness to use its Python. + +## SEC identity + +SEC requests require a real contact identity under the SEC fair-access policy: ```bash -npx skills add eggmasonvalue/sec-edgar-skill +export EDGAR_IDENTITY="Jane Analyst jane@example.com" ``` -## How it works +Do not commit it. The 13f.info convenience queries and local-file utilities do not require an +SEC identity. Verify the full environment with: -Filings are huge, so the skill keeps them on disk and pulls only what's needed into the -agent's context: +```bash +uv run python scripts/test_setup.py --live +``` -1. **Orient** with `scripts/orient.py` (company summary + filing-mix survey) to decide what to fetch. -2. **Download** filings to a local cache (`./sec-cache/{TICKER}/`) as clean Markdown. -3. **Map** a large filing to a heading→line table of contents. -4. **Search** the cache with native grep and read only the matching line ranges. +## Retrieval model -The cache location is configurable (`$SEC_CACHE_DIR` or `--cache-dir`) and filenames are -deterministic (keyed by SEC accession number), so re-runs reuse cached files instead of -re-downloading. +1. Survey a company's filing mix when the relevant form is unknown. +2. Select a filing by company/form/period or exact SEC accession. +3. Save filing text as Markdown and statements as CSV. +4. Search the local cache and read only relevant ranges. +5. Use structured item or XBRL extraction where the filing supports it. + +Accession-keyed filing artifacts are reused unless `--force` is passed. Rolling ownership +summaries refresh because their source set can change. ## Data sources -Filing data comes from the public SEC EDGAR system via the open-source `edgartools` -library. Respect the source's terms and the SEC fair-access policy — which is why a -contact identity is required. +Filing data comes from SEC EDGAR through the open-source `edgartools` library. Routine 13F +holder, manager, and position-history queries are distilled through 13f.info; generated reports +surface the underlying SEC reporting period, manager CIK, and accession rather than provider +navigation links. Raw EDGAR remains the verification and deep-field route. --- -Part of the [SecStack skills](../README.md) collection. + +Part of the [SecStack skills](https://github.com/eggmasonvalue/secstack) collection. diff --git a/skills/sec-edgar-skill/SKILL.md b/skills/sec-edgar-skill/SKILL.md index 2657629..5971374 100644 --- a/skills/sec-edgar-skill/SKILL.md +++ b/skills/sec-edgar-skill/SKILL.md @@ -1,171 +1,149 @@ --- name: sec-edgar-skill description: >- - Retrieve and extract SEC EDGAR filings and ownership data for US-listed companies. Use this - whenever a task touches a company's filings, financials, ownership, governance, or - institutional holders — even if EDGAR is not named explicitly. Covers 10-K/10-Q/8-K, - 20-F/6-K foreign filings, XBRL financial statements, insider transactions, 13F institutional - holdings (via 13f.info), 13D/13G blockholdings, and DEF 14A proxy/compensation. Always start - by running scripts/orient.py, then pull only the sections you need. Unopinionated data layer: - it fetches and extracts token-efficiently; it does not decide what is significant. + Retrieve, verify, and inspect SEC filings and filing-derived financial, ownership, + compensation, and governance data. Use when current or auditable evidence must be pulled + from EDGAR: 10-K/10-Q/8-K, 20-F/40-F/6-K, XBRL statements, Form 4 insider transactions, + 13F institutional holdings, 13D/13G blockholders, or proxy disclosures. The generic filing + tools also accept other SEC forms. Do not trigger merely because an analysis mentions + financials; use it when source material needs to be obtained or checked. --- # SEC EDGAR Research Skill -A toolkit for retrieving and extracting SEC EDGAR filings for US-listed companies — -efficiently, within a token budget, using `edgartools`. - -## What this skill is — and is not - -This is the **tools layer** of a research stack. It knows *how* to find, download, and -extract SEC filing data, plus the library mechanics to do it reliably. It is -deliberately **unopinionated**: it does not judge what is a good number, a red flag, or -worth looking at. That judgment belongs to whatever **analytical-framework skill** is -driving (e.g. a value-investing framework); presentation belongs to a downstream -consumer skill. Keep this layer neutral so any framework can compose on top of it. - -Provide capability and facts; let the caller reason. The only hard requirements here are -mechanical (set an SEC identity, or requests are blocked) — never analytical. - -## Setup - -Run `python scripts/test_setup.py --live` to verify dependencies and SEC identity. -If identity is missing, the error message tells the user exactly how to set -`$EDGAR_IDENTITY` (see the [profile setup](../../README.md#one-time-runtime-setup) for the full -explanation). Every script reads it automatically; you can also pass `--identity` -per-call. - -The scripts handle environment hazards on import (UTF-8 stdout on Windows, truststore -for corporate proxies). When writing **inline Python** that calls `edgartools` directly -(not via a bundled script), replicate that preamble — see `references/guide_core.md` -§ "Inline Python preamble." - -## How to work efficiently: pull only what you need - -Filings are huge — a 10-K can exceed 100k words. Loading one whole wastes the context -window and buries the signal. The method is to keep documents **on disk** and pull only -the exact lines you need into context. Four phases: - -1. **Orient first — always run `scripts/orient.py`.** `python scripts/orient.py --ticker ` - prints the company's `.to_context()` summary, surveys the **mix of forms it has actually - filed** (with date ranges), and lists the most recent filings — the cheapest way to see what - a company files *now* and how that has changed, so you fetch the right forms instead of - assuming a form set. This is the **non-negotiable first step — run it before any web - search, even for breaking news.** When the user says "just reported" or "a few hours ago," - orient.py will show the 8-K filed today immediately; then fetch it with - `fetch_filing.py --form 8-K --date --attachment list` to get the press-release - exhibit (Exhibit 99.1). The filing is always faster and more authoritative than a web - search for what the company itself disclosed. (For finer control, `.to_context()` - on a `Company`, filing collection, or `XBRL` object gives the same preview inline — see - `guide_core.md`.) -2. **Download to the cache as Markdown.** Use the scripts to write filings to disk as - clean Markdown — `edgartools` converts SEC HTML, stripping layout bloat to roughly a - tenth of the size, and the result is greppable. -3. **Get just the section you need.** Item-addressable forms (10-K/10-Q/8-K/20-F) let you - list a filing's SEC item codes and pull one section by code — no scanning. For a full - report, `scripts/list_headings.py` maps its `#` headers to line numbers; tabular - free-form filings (e.g. DEF 14A) instead carry their own table of contents up top — - read it, then grep. `guide_filings.md` has the mechanics. -4. **Search, then read precisely.** Use your native grep/ripgrep over the cached files - to find the lines that matter, then read just those ranges. Let grep and disk do the - heavy lifting; spend context only on the paragraphs you actually need. - -## The cache - -Downloads go to `//
__[__].md` -(statements as `…__.csv`). The root resolves as `--cache-dir` > -`$SEC_CACHE_DIR` > `./sec-cache` — workspace-relative so your grep tool finds it by -default, and persistent across runs so you don't re-hit the SEC. Filenames are -deterministic and keyed by the globally-unique accession number, so **before -downloading, check whether the file already exists** (list or glob `//`) -and reuse it. Every script prints the absolute path(s) it wrote to stdout. - -> If your grep tool uses ripgrep and the cache is gitignored, ripgrep skips it by -> default. Either point the search at `//` explicitly, or pass -> `--no-ignore`. Resolve the path from a script's stdout or `$SEC_CACHE_DIR` — never -> hard-code an absolute cache path. - -## Reference guides (read the relevant one before extracting) - -Each guide is loaded only when its domain is in play, so you carry just the rules you -need. Read the matching guide first — it holds the item codes, taxonomies, and library -quirks that make extraction correct. +Retrieve filing-derived evidence without loading entire filings into context. -| Guide | Use it for | -| :-- | :-- | -| `references/guide_core.md` | The mechanics behind `scripts/orient.py`: resolving a company, listing/filtering filings, surveying the filing mix, `.to_context()` previews, and `.docs` self-help. Read it to drive orientation inline or go beyond the script. | -| `references/guide_filings.md` | Filing text: pulling a section by SEC item code (10-K/10-Q/8-K/20-F) vs. navigating free-form filings (DEF 14A/6-K) by their own contents, plus attachments and exhibits (incl. 6-K Exhibit 99.1). | -| `references/guide_financials.md` | XBRL financial statements and individual facts (US-GAAP and IFRS), and the period-aggregation pitfalls. | -| `references/guide_ownership.md` | Insider transactions (Forms 3/4/5), beneficial ownership, and executive compensation (DEF 14A; Form 20-F Item 6 for foreign issuers). For the common case — “what are insiders buying/selling?” — use `scripts/fetch_insider_trades.py` directly; no guide needed. | -| `references/guide_holdings.md` | **Deep route only:** raw 13F via edgartools (voting authority, amendments, specific holdings) and 5%+ blockholders (13D/13G). For the common case — “who owns this stock?” — use `scripts/fetch_13f_holders.py` directly (see Scripts above); no guide needed. | +## Role in the research stack + +Prefer an SEC filing over model memory or a secondary summary for facts the filing discloses. +When an analytical skill is driving the task, provide the requested evidence and provenance +without substituting an investment conclusion. Use company investor relations, another +regulator, or reputable web sources when the information is not yet on EDGAR or is outside +EDGAR's scope. -## Scripts +## Runtime and paths -Run them with the project's Python. Each prints the absolute cache path(s) it wrote to -stdout and logs progress to stderr. **`--help` is the authoritative flag reference** — -the list below shows one canonical invocation each: +Resolve bundled paths relative to this `SKILL.md`. Invoke scripts by their resolved absolute +path while keeping the shell working directory at the research workspace; this keeps the +default `./sec-cache` beside the work rather than inside the installed skill. Alternatively, +pass `--cache-dir`. + +SEC requests require a real contact identity: ```bash -# Orient first: company summary + filing-mix survey + recent filings -python scripts/orient.py --ticker AAPL +export EDGAR_IDENTITY="Jane Analyst jane@example.com" +``` -# One filing (full), a single section, or an attachment — into the cache -python scripts/fetch_filing.py --ticker AAPL --form 10-K --year 2023 -python scripts/fetch_filing.py --ticker AAPL --form 10-K --year 2023 --section "Item 1A" # or: --section list -python scripts/fetch_filing.py --ticker WIX --form 6-K --attachment "ex-99.1" # or: list | all | +Every SEC-facing script also accepts `--identity`. The 13f.info convenience script and local +`list_headings.py` do not need an SEC identity. Diagnose the environment with: -# Target a filing by date (e.g. an 8-K filed today) instead of just --year: -python scripts/fetch_filing.py --ticker AAPL --form 8-K --date 2026-06-15 +```bash +python "/scripts/test_setup.py" --live +``` -# Many filings across a year range (add --attachments to capture e.g. 6-K exhibits) -python scripts/fetch_filings.py --ticker AAPL --form 10-Q --start-year 2022 --end-year 2024 +## Choose the shortest reliable route + +1. **Orient when discovery is needed.** Run `orient.py` when the relevant form or accession is + unknown. It prints company identity, the recent filing mix, and recent accessions without + downloading filing contents. Skip orientation when the user supplied an accession or + cached file, or when the task is solely a 13f.info convenience query. +2. **Select exactly.** Use company/form/period selectors for ordinary retrieval. Use + `--accession` when the user supplied one or a candidate-listing command identified one. + `--on-or-before` is deliberately explicit: it may select an earlier filing. +3. **Keep large documents on disk.** Download Markdown or CSV into the cache. Search cached + Markdown with native grep/ripgrep, then read only relevant line ranges. +4. **Prefer structured routes where they are sound.** Pull item-addressable sections and XBRL + statements directly. For free-form filings, inspect their contents and attachments. +5. **Treat failure as failure.** Do not turn a network or parser error into a factual “no data” + conclusion. Use the reported candidates, accession route, relevant guide, or edgartools + `.docs` to recover. + +For very recent disclosures, check EDGAR promptly but do not assume a filing must already +exist. A company release can precede its SEC filing. + +## Cache behavior + +The root resolves as `--cache-dir` > `$SEC_CACHE_DIR` > `./sec-cache`. Filing artifacts use: + +```text +/__[__].md +/____.csv +``` -# XBRL statements (income | balance | cashflow | all) -> CSV (annual or quarterly) -python scripts/parse_financials.py --ticker AAPL --year 2023 --statement all -python scripts/parse_financials.py --ticker AAPL --year 2024 --quarter 1 --statement all +Accession-keyed filing artifacts are immutable enough to reuse by default; pass `--force` to +regenerate them. Rolling insider and 13F summaries refresh because late filings or amendments +can change their source set. If ripgrep skips a gitignored cache, point it directly at the +company directory or use `--no-ignore`. -# Table of contents for a large cached filing -python scripts/list_headings.py --file sec-cache/AAPL/10-K_2023-11-03_0000320193-23-000106.md +## Reference guides -# Insider transactions — what are insiders buying/selling? (Form 4) -python scripts/fetch_insider_trades.py --ticker AAPL -python scripts/fetch_insider_trades.py --ticker AAPL --start 2025-01-01 --end 2026-06-17 -python scripts/fetch_insider_trades.py --ticker AAPL --start 2025-06-01 --buys-only +Read only the guide needed for the question: -# 13F institutional holders — who owns this stock? (via 13f.info, no SEC identity needed) -python scripts/fetch_13f_holders.py --ticker AAPL --top 15 +| Guide | Use it for | +| :-- | :-- | +| `references/guide_core.md` | Company resolution, filing discovery, `.to_context()`, exact accessions, and `.docs`. | +| `references/guide_filings.md` | Full filing text, SEC item codes, free-form navigation, and exhibits. | +| `references/guide_financials.md` | XBRL statements/facts, period semantics, and noisy financial 6-Ks. | +| `references/guide_ownership.md` | Forms 3/4/5, insider transactions, and Section 16 limitations. | +| `references/guide_proxy.md` | Beneficial ownership, board/governance, compensation, related parties, auditors, dilution, proposals, and voting. | +| `references/guide_holdings.md` | 13F institutional holdings and 13D/13G blockholder schedules. | -# 13F holder history — how has institutional ownership changed? -python scripts/fetch_13f_holders.py --ticker AAPL --history +## Script routes -# 13F manager search — what does a specific fund hold? -python scripts/fetch_13f_holders.py --manager "Berkshire Hathaway" +`--help` is the authoritative flag reference. Artifact-producing commands emit absolute paths +to stdout and progress to stderr; orientation, diagnostics, and list modes print reports. -# 13F cross-reference — one manager's position history in one stock -python scripts/fetch_13f_holders.py --cik 0000906304 --cusip 205826209 +```bash +# Discover forms and recent accessions +python "/scripts/orient.py" --ticker AAPL -# Environment diagnostics -python scripts/test_setup.py --live -``` +# Filing by ordinary selectors or exact accession +python "/scripts/fetch_filing.py" --ticker AAPL --form 10-K --year 2025 +python "/scripts/fetch_filing.py" --accession 0000320193-25-000079 -## When the API surprises you: self-heal with `.docs` +# One item or the filing's actual item list +python "/scripts/fetch_filing.py" --ticker AAPL --form 10-K --year 2025 --section "Item 1A" +python "/scripts/fetch_filing.py" --accession 0000320193-25-000079 --section list -`edgartools` documents itself at runtime. If a method or attribute isn't what you -expected, query it inline instead of guessing — this recovers from most API uncertainty -without leaving the session: +# An exhibit, or the nearest matching filing on/before a date +python "/scripts/fetch_filing.py" --ticker WIX --form 6-K --attachment "ex-99.1" +python "/scripts/fetch_filing.py" --ticker AAPL --form 8-K --on-or-before 2026-06-15 -```python -company.docs # full API guide for the object -company.docs.search("xbrl") # search it for a topic -``` +# Bulk archive: intentionally includes originals and amendments +python "/scripts/fetch_filings.py" --ticker AAPL --form 10-Q --start-year 2022 --end-year 2024 -A tool error is almost always a fixable usage detail, not a dead end. When a script or call -fails, **recover here** — re-run `scripts/orient.py`, query `.docs`, or read the relevant -guide — rather than abandoning EDGAR for web search. The filings are the authoritative, -auditable source; don't let a transient error push the work onto unverifiable web results. +# Structured statements; period selection never falls back to a different report +python "/scripts/parse_financials.py" --ticker AAPL --year 2025 --statement all +python "/scripts/parse_financials.py" --accession 0000320193-25-000079 --statement all + +# Local heading map +python "/scripts/list_headings.py" --file sec-cache/AAPL/.md + +# Form 4 transactions +python "/scripts/fetch_insider_trades.py" --ticker AAPL +python "/scripts/fetch_insider_trades.py" --ticker AAPL --start 2025-06-01 --buys-only + +# Distilled 13F queries +python "/scripts/fetch_13f_holders.py" --ticker AAPL +python "/scripts/fetch_13f_holders.py" --ticker AAPL --history +python "/scripts/fetch_13f_holders.py" --manager "Berkshire Hathaway" +python "/scripts/fetch_13f_holders.py" --manager 0001067983 --ticker AAPL +``` -**Amendment vs. original filing.** `fetch_filing.py` always skips amended forms -(10-K/A, 10-Q/A, etc.) and picks the most recent *original* filing. Amendments typically -contain only the amended items (e.g. Part III), not the full filing, so silently picking -one would lose most of the content. If you specifically need an amendment's items, fetch -it by accession number using inline Python as shown in `guide_financials.md`. +The 13F convenience script uses 13f.info to resolve and distill routine holdings queries but +its reports expose the underlying SEC period, CIK, and accession rather than the intermediary. +Use raw EDGAR only for fields the distilled route does not provide—such as voting authority or +investment discretion—or to verify missing or inconsistent data. + +## Important filing semantics + +- `fetch_filing.py` selects original forms for ordinary company/form requests because an + amendment may contain only changed items. Exact `--accession` retrieval can fetch either. +- `fetch_filings.py` intentionally preserves both originals and amendments in a bulk archive. +- A 6-K is not a standardized quarterly report. `parse_financials.py` prefers a matching 10-Q; + otherwise it classifies financial 6-K candidates. If several contain statements, it lists + their accessions instead of guessing. +- Form 4 reports distinguish transaction date from filing date and expose partial parsing. +- A successful generated report is complete for its advertised fields. Do not browse the + intermediary provider after a successful 13F query; use the underlying SEC accession when + deeper verification is required. diff --git a/skills/sec-edgar-skill/references/guide_core.md b/skills/sec-edgar-skill/references/guide_core.md index 809e666..2324a86 100644 --- a/skills/sec-edgar-skill/references/guide_core.md +++ b/skills/sec-edgar-skill/references/guide_core.md @@ -1,10 +1,9 @@ # Core — company lookup, filing discovery, and self-help -Start here. This guide covers resolving a company, listing and filtering its filings, -and the two built-in efficiency tools — `.to_context()` previews and the `.docs` -self-help system. `scripts/orient.py` wraps the orientation steps below into one command; -this guide is the mechanics behind it, for when you drive them inline. The other guides -build on these basics. +Use this guide when the relevant form or accession is not already known, or when a +bundled script does not expose a needed filing filter. `scripts/orient.py` wraps company +identity, filing-mix, and recent-accession discovery into one command. Skip orientation for +an exact accession or an already-cached local file. ## Resolve a company @@ -29,7 +28,7 @@ tickers. ```python filings = company.get_filings() # everything -filings = company.get_filings(form="10-Q", year=2024) # by form + year +filings = company.get_filings(form="10-Q", year=2024, amendments=False) filings = company.get_filings(quarter=4, year=2024) # by quarter filings = company.get_filings(date="2023-01-01:2023-12-31") # by date range ``` @@ -37,16 +36,29 @@ filings = company.get_filings(date="2023-01-01:2023-12-31") # by date range Collections support indexing, slicing, and `.latest()`: ```python -latest_10k = company.get_filings(form="10-K").latest() +latest_10k = company.get_filings(form="10-K", amendments=False).latest() recent = filings[0:10] ``` +## Resolve an exact accession + +When orientation or another command reports an accession, use it directly rather than +reconstructing the filing from date and form: + +```python +from edgar import find + +filing = find("0000320193-25-000079") +``` + +The bundled filing and financial scripts expose the same route as `--accession`. Do not +combine an accession with company/period selectors; the accession is already globally unique. + ## Survey a company's filing mix `scripts/orient.py` does this for you (with per-form date ranges and the most recent filings); reach for it first. To do it inline — or to tabulate a custom window — pull the -collection to a DataFrame. This is a neutral mechanic; which forms are relevant to your -question is for you (or the framework driving you) to decide. +collection to a DataFrame. Which forms are relevant depends on the requested evidence. ```python df = company.get_filings(date="2024-01-01:2025-12-31").to_pandas() diff --git a/skills/sec-edgar-skill/references/guide_filings.md b/skills/sec-edgar-skill/references/guide_filings.md index fb85eda..bf5caba 100644 --- a/skills/sec-edgar-skill/references/guide_filings.md +++ b/skills/sec-edgar-skill/references/guide_filings.md @@ -20,7 +20,7 @@ Periodic and current reports parse into a typed object whose sections are keyed item code. You don't discover the structure — you ask for the item directly. ```python -filing = company.get_filings(form="10-K").latest() +filing = company.get_filings(form="10-K", amendments=False).latest() report = filing.obj() # TenK / TenQ / CurrentReport / TwentyF / ... report.items # -> ['Item 1', 'Item 1A', 'Item 1B', ...] actually present risk = report["Item 1A"] # just that item's text, not the whole 100k-word filing @@ -94,7 +94,9 @@ text = attachments[1].markdown() # convert one exhibit to Markdown ``` Script: `fetch_filing.py --attachment "ex-99.1"` (or `list` | `all` | an index), or -`fetch_filings.py --attachments` to capture exhibits across a whole year range. +`fetch_filings.py --attachments` to capture exhibits across a whole year range. When an +accession is known, combine `--accession` with the attachment selector so no date/form guess +is involved. > **Index attachments via a list.** `filing.attachments` looks items up by their 1-based > SEC *sequence number*, which can skip values — so integer indexing on the raw collection @@ -108,8 +110,6 @@ Script: `fetch_filing.py --attachment "ex-99.1"` (or `list` | `all` | an index), ## Then search locally -Once a filing (or section, or exhibit) is in the cache, search it with your native -grep/ripgrep and read the matching line ranges. Don't run text searches through the library -— that's slower and round-trips to remote endpoints. *What* you search for, and what you -make of it, is yours (or your framework's) to decide; this skill just makes the text fast to -reach. +Once a filing, section, or exhibit is in the cache, search it with native grep/ripgrep and +read the matching line ranges. Avoid remote text searches when the local Markdown already +contains the source text. diff --git a/skills/sec-edgar-skill/references/guide_financials.md b/skills/sec-edgar-skill/references/guide_financials.md index e3c562c..0b85a91 100644 --- a/skills/sec-edgar-skill/references/guide_financials.md +++ b/skills/sec-edgar-skill/references/guide_financials.md @@ -1,65 +1,103 @@ -# Financials — XBRL statements and facts +# Financials — XBRL statements, facts, and periods -How to pull structured financial statements (Income, Balance Sheet, Cash Flow) and -individual XBRL facts, for both US-GAAP and IFRS filers. `scripts/parse_financials.py` -wraps statement extraction to CSV. +Use XBRL for structured income statements, balance sheets, cash-flow statements, and +individual facts. `scripts/parse_financials.py` writes statements to accession-keyed CSVs. +It never substitutes a different filing when the requested filing period has no match. -## Get filings for parsing +## Select the filing + +For ordinary annual or domestic quarterly reports: ```python -filings = company.get_filings(form=["10-K", "20-F", "40-F"], year=2024, amendments=False) +filings = company.get_filings( + form=["10-K", "20-F", "40-F"], + year=2025, + amendments=False, +) filing = filings.latest() ``` -> **Pass `year` to `get_filings`, not `.filter()`.** `EntityFilings.filter(year=...)` -> raises `TypeError` — `filter` doesn't accept `year`. -> -> **Use `amendments=False`.** Amendments (`10-K/A`, etc.) often carry only minor text -> changes and lack complete XBRL statement trees; exclude them to get the primary -> statements. +Pass `year` to `get_filings`, not `EntityFilings.filter()`, which does not accept it. Exclude +amendments for ordinary statement extraction because an amendment may lack the complete XBRL +statement tree. + +When the exact record is known, prefer its accession: + +```python +from edgar import find + +filing = find("0000320193-25-000079") +``` + +The script equivalents are: + +```bash +python scripts/parse_financials.py --ticker AAPL --year 2025 --statement all +python scripts/parse_financials.py --accession 0000320193-25-000079 --statement all +``` + +A missing period is an error. Read the nearby candidates in stderr and rerun deliberately; +do not silently use the latest report. + +## Financial 6-Ks require candidate selection + +A 6-K is a foreign private issuer's general current report, not a standardized quarterly +report. It can contain financing, governance, operational, or other disclosures. Some 6-Ks +carry tagged annual, half-year, nine-month, revised, or voluntary interim financial statements. + +For a requested filing window, `parse_financials.py`: + +1. prefers a matching original 10-Q when one exists; +2. otherwise surveys 6-K and 6-K/A filings in the window; +3. looks for XBRL metadata, statement objects, and financial attachment descriptions; +4. extracts when exactly one filing has parseable statements; +5. lists candidate accessions instead of guessing when several match; +6. identifies likely textual financial 6-Ks when no statements can be parsed. + +Use the reported accession for a second, exact call. For a textual candidate, fetch that +accession with `fetch_filing.py --attachment list` and pull the relevant exhibit. -Annual reports (10-K / 20-F / 40-F) and quarterly reports (10-Q / 6-K) carry the -complete statement trees and can be parsed to CSV using `parse_financials.py`. +XBRL proves that tagged information exists; it does not by itself prove that the filing is a +calendar-quarter report. Verify statement period ends and durations before labeling the data. -## Parse statements from a filing +## Parse statements ```python xbrl = filing.xbrl() -print(xbrl.to_context()) # lists available statements +print(xbrl.to_context()) income = xbrl.statements.income_statement() balance = xbrl.statements.balance_sheet() -cash = xbrl.statements.cashflow_statement() # note: "cashflow", no underscore +cash = xbrl.statements.cashflow_statement() df = income.to_dataframe() ``` -> Statement accessors live on `xbrl.statements`, not on the `XBRL` object itself, and the -> cash-flow method is `cashflow_statement()` (no underscore in "cashflow"). +Statement accessors live on `xbrl.statements`; the cash-flow accessor is +`cashflow_statement()` without an underscore inside “cashflow.” -## Multi-period history from the company +## Multi-period company financials ```python -fin = company.get_financials() -print(fin.to_context()) -df = fin.income_statement().to_dataframe() # methods are directly on this object +financials = company.get_financials() +print(financials.to_context()) +df = financials.income_statement().to_dataframe() ``` -On the object returned by `get_financials()`, the statement methods are direct — not -under `.statements`. +On the object returned by `get_financials()`, statement methods are direct rather than under +`.statements`. -> **Don't blindly `.mean()` / `.sum()` across periods.** A multi-period pull mixes -> current-period values with prior-period comparatives. For balance-sheet (instant) -> facts, filter `period_instant == report_date`; for income/cash-flow (duration) facts, -> filter `period_end == report_date` and sanity-check the duration (~90 days quarterly, -> ~360 days annual). Otherwise you average current figures with comparatives and corrupt -> the series. +Do not blindly average or sum rows across periods. A pull can mix current values with prior +comparatives. For balance-sheet instant facts, filter to the intended `period_instant`; for +income and cash-flow duration facts, verify `period_end` and duration. Quarterly, half-year, +nine-month, and annual durations are not interchangeable. ## Individual facts — US-GAAP and IFRS ```python -rev_us = xbrl.get_fact("us-gaap:Revenues") -rev_ifrs = xbrl.get_fact("ifrs-full:Revenue") # foreign issuers often file IFRS +revenue_us = xbrl.get_fact("us-gaap:Revenues") +revenue_ifrs = xbrl.get_fact("ifrs-full:Revenue") ``` -If US-GAAP tags come back empty for a foreign private issuer, try the IFRS equivalent — -20-F filers frequently report under IFRS rather than US-GAAP. +Foreign private issuers frequently report under IFRS. If a US-GAAP concept is absent, inspect +the filing taxonomy and use the corresponding IFRS or issuer-extension concept rather than +assuming the fact is missing. diff --git a/skills/sec-edgar-skill/references/guide_holdings.md b/skills/sec-edgar-skill/references/guide_holdings.md index 54b2f8b..ad28ef3 100644 --- a/skills/sec-edgar-skill/references/guide_holdings.md +++ b/skills/sec-edgar-skill/references/guide_holdings.md @@ -1,49 +1,79 @@ -# Holdings — deep route (edgartools) and blockholders (13D/13G) +# Institutional holdings and blockholders -> **For the common case** — "who owns this stock?", "how has ownership changed?", -> "what does this fund hold?" — use `scripts/fetch_13f_holders.py` directly. -> It is already listed in the Scripts section of SKILL.md and needs no guide. -> Read this file only when you need **raw 13F filing data** (voting authority, -> amendments, specific holding details) or **blockholder schedules** (13D/13G). +Use the distilled 13F script for routine stock-, manager-, and manager-position questions. Use +raw EDGAR when a required filing field is not surfaced, the distilled result is missing or +inconsistent, or the user asks for source-level verification. + +## Distilled 13F routes + +```bash +# Which reporting managers disclosed this security? +python scripts/fetch_13f_holders.py --ticker AAPL + +# Aggregate reporting-manager/share history +python scripts/fetch_13f_holders.py --ticker AAPL --history + +# Latest complete disclosed portfolio for a manager +python scripts/fetch_13f_holders.py --manager "Berkshire Hathaway" + +# Latest quarter-over-quarter share changes plus filing history +python scripts/fetch_13f_holders.py --manager "Berkshire Hathaway" --history + +# One manager's disclosed position history in one security +python scripts/fetch_13f_holders.py --manager 0001067983 --ticker AAPL +``` + +The script resolves manager CIKs and security CUSIPs internally. Reports expose underlying SEC +periods and identifiers rather than intermediary-provider links. A successful report is the +complete normal surface; do not browse the provider looking for omitted routine fields. + +13F is delayed and incomplete as a picture of ownership. It covers reportable long positions +of filing managers, not retail holders, all institutions, short positions, or every security. +Treat “who owns this stock?” as shorthand for “which Form 13F managers disclosed this CUSIP?” +Reconcile multiple share classes and CUSIPs before aggregating. ## Raw 13F via edgartools -13F is filed by the **investment manager**, not the operating company. Query the manager: +A 13F is filed by the investment manager, not the operating company: ```python manager = Company("Magnetar Capital LLC") -f13 = manager.get_filings(form="13F-HR").latest().obj() -df = f13.holdings # the holdings DataFrame +filing = manager.get_filings(form="13F-HR", amendments=False).latest() +thirteen_f = filing.obj() +holdings = thirteen_f.holdings ``` -> Operating companies (AAPL, etc.) do **not** file 13F — querying their CIK for `13F-HR` -> returns an empty collection. You must query the fund manager's name or CIK. -> -> Holdings live on the `.holdings` attribute; the 13F object has no `.to_dataframe()`. +Use raw holdings for voting authority, investment discretion, other-manager fields, precise +amendment treatment, or verification. Holdings are already a DataFrame; the 13F object does +not need `.to_dataframe()`. -Filter the holdings DataFrame by CUSIP or issuer name to confirm whether a given manager -holds a target stock (you query each manager you care about). +An operating company's CIK normally does not file the 13F portfolios that hold its stock. A +stock-centric reverse lookup therefore belongs to the distilled route. -## Blockholders (Schedules 13D / 13G) +## Schedules 13D and 13G -13D signals an active/control intent; 13G signals a passive stake. Query both, and -include the `"SC "`-prefixed names — EDGAR often indexes these schedules that way, so -omitting them returns empty results: +Schedules 13D and 13G disclose beneficial ownership above the applicable threshold under +different eligibility and reporting regimes. Do not reduce the distinction to “active” versus +“passive”: Schedule 13G also covers qualified institutional and exempt investors. + +Query the schedule names used by EDGAR, including `SC` prefixes and amendments: ```python -blocks = company.get_filings(form=["13D", "13G", "SC 13D", "SC 13G", "SC 13D/A", "SC 13G/A"]) -sched = blocks.latest().obj() # Schedule13D / Schedule13G structured object -sched.reporting_persons # who holds — each with voting / dispositive power -sched.issuer_info # the subject company (name, CIK, CUSIP) -sched.total_percent # aggregate % of class -sched.items.item4_purpose_of_transaction # 13D Item 4 narrative, when present +schedules = company.get_filings(form=["13D", "13G", "SC 13D", "SC 13G", "SC 13D/A", "SC 13G/A"]) +schedule = schedules.latest().obj() +schedule.reporting_persons +schedule.issuer_info +schedule.total_percent +schedule.items.item4_purpose_of_transaction ``` -These schedules parse into a **structured object**, not an item-addressable one — there is -no `sched["Item 4"]`. The per-item values live on `sched.items` (a dataclass with fields -like `item4_purpose_of_transaction` and `item5_percentage_of_class`). +These parse into structured schedule objects rather than item-addressable company reports. +Fields such as the 13D purpose narrative live on `schedule.items`; there is no generic +`schedule["Item 4"]` contract. + +Follow the amendment chain for a reporting person instead of treating every accession as a +separate current holder. Preserve reporting-person identity, share class, measurement date, +shares, percentage, voting/dispositive power, and accession. -> **Older filings lack structured tags.** Before EDGAR's late-2024 structured-XML mandate, -> `sched.has_structured_data` is `False` and the parsed fields come back `None`/`0`. For -> those, save the full Markdown (`fetch_filing.py`, no `--section`) and grep the ownership -> figure and the Item 4 purpose out of the text instead. +Older schedules may lack structured XML. When parsed fields are absent, fetch the exact +accession as full Markdown and inspect the ownership tables and Item 4 narrative directly. diff --git a/skills/sec-edgar-skill/references/guide_ownership.md b/skills/sec-edgar-skill/references/guide_ownership.md index 00b8616..a15c77c 100644 --- a/skills/sec-edgar-skill/references/guide_ownership.md +++ b/skills/sec-edgar-skill/references/guide_ownership.md @@ -1,64 +1,63 @@ -# Ownership & compensation — insiders, proxies, and exec pay +# Insider ownership — Forms 3, 4, and 5 -Two related questions about the people who run and own a company: what insiders are -buying and selling (Forms 3/4/5), and how executives are paid (the DEF 14A proxy, or -Form 20-F Item 6 for foreign issuers). For *external* 5%+ holders and institutions, see -`guide_holdings.md` instead. +Use this guide for Section 16 ownership reports by directors, officers, and 10% holders. For +annual-meeting beneficial ownership, governance, and compensation, read `guide_proxy.md`. For +external blockholders and institutional managers, read `guide_holdings.md`. -## Insider transactions (Forms 3, 4, 5) +## What the forms mean -There is no `company.get_insiders()`. Query the ownership forms directly: +- **Form 3:** initial beneficial-ownership statement when a person becomes subject to Section 16. +- **Form 4:** changes in ownership, generally reported within two business days. +- **Form 5:** selected annual transactions that were exempt or reported late. + +There is no `company.get_insiders()` requirement for retrieval; query the forms directly: ```python -form4s = company.get_filings(form="4") # 3 = initial, 4 = changes, 5 = annual -latest = form4s.latest().obj() # parse the XML into a structured object +form4s = company.get_filings(form="4", amendments=False) +latest = form4s.latest().obj() print(latest.insider_name, latest.position) -df = latest.to_dataframe() +dataframe = latest.to_dataframe() ``` -Access individual trades via DataFrames or the activities helper — not by iterating -`non_derivative_transactions` (which isn't exposed as a standard list): - -```python -df_market = latest.market_trades # open-market buys / sells -df_options = latest.option_exercises -for act in latest.get_transaction_activities(): - print(act.transaction_type, act.code, act.shares, act.price_per_share) - # transaction codes: P = purchase, S = sale, M = option exercise, F = tax withholding -``` +The bundled `fetch_insider_trades.py` is the normal company-level activity route. It queries +original Form 4s to avoid counting an original and amendment twice, distinguishes transaction +date from filing date, and reports partial parser coverage. Inspect a Form 4/A directly when a +specific correction matters. -## Executive compensation — DEF 14A (domestic filers) +## Transaction data -US filers disclose executive compensation in the annual proxy (DEF 14A) and incorporate it -by reference into the 10-K rather than printing it there — so look in the proxy, not the -10-K. A proxy is a free-form filing: `filing.obj()` is a `ProxyStatement` with no item -codes, so you can't pull a section by code the way you can from a 10-K. Download it, read -the table of contents in its first ~100 lines to see its sections, then grep the body for -the part you want. +Use the DataFrame or activity helper rather than assuming internal transaction collections are +ordinary lists: ```python -proxy = company.get_filings(form="DEF 14A").latest() -text = proxy.markdown() # save to the cache, then read its ToC + grep +dataframe = latest.market_trades +option_exercises = latest.option_exercises +for activity in latest.get_transaction_activities(): + print( + activity.transaction_type, + activity.code, + activity.shares, + activity.price_per_share, + ) ``` -The compensation disclosures (the Summary Compensation Table and the rest) are mandated by -Regulation S-K Item 402 — so a proxy's contents are predictable from that rule. Which tables -and narrative you pull, and what you make of them, is the caller's call. See -`guide_filings.md` for the item-addressable-vs-free-form distinction in general. +Common transaction codes include: -## Foreign private issuers +- `P` — open-market purchase; +- `S` — open-market sale; +- `M` — option exercise; +- `F` — tax withholding; +- `A` — award or grant; +- `G` — gift. + +Only code `P` is an open-market insider purchase. Preserve the transaction date, filing date, +accession, shares, price, and post-transaction holdings so the record remains auditable. -FPIs are exempt from several US ownership and governance rules, which changes *where* the -data is: +## Foreign private issuers -- **Section 16 exemption.** FPIs (and Canadian MJDS filers) don't file Forms 3/4/5, so - insider-transaction data won't appear on EDGAR. Check the home-jurisdiction regulator - (e.g. SEDAR+ for Canada) instead. -- **No DEF 14A.** FPIs don't file US proxies. Compensation is in **Form 20-F Item 6.B** - and share ownership in **Item 6.E**. Many FPIs disclose pay only in aggregate unless - their home-country rules require individual figures. +Foreign private issuers, including Canadian MJDS filers, are generally exempt from Section 16 +Forms 3/4/5. Absence of Form 4 filings is therefore not evidence that no insider transaction +occurred. Check the home-jurisdiction regulator when the task requires that information. -> **20-F Item 7.A boundary trap.** When slicing "Item 7.A Major Shareholders" by text, -> note the full Item 7 title is "Major Shareholders **and Related Party Transactions**." -> Using "Related Party Transactions" as the end boundary truncates early, because it also -> appears in the title — terminate the slice on "Item 8" instead. +Foreign-issuer compensation and management ownership are generally found in Form 20-F Item 6; +see `guide_proxy.md` for the disclosure map. diff --git a/skills/sec-edgar-skill/references/guide_proxy.md b/skills/sec-edgar-skill/references/guide_proxy.md new file mode 100644 index 0000000..3d3eccf --- /dev/null +++ b/skills/sec-edgar-skill/references/guide_proxy.md @@ -0,0 +1,107 @@ +# Proxy and governance disclosures + +Use this guide for beneficial ownership, board structure, compensation, related parties, +auditors, equity plans, proposals, and voting. Domestic issuers generally disclose these in +DEF 14A; foreign private issuers generally use Form 20-F and home-jurisdiction materials. + +## Navigate a domestic proxy + +A DEF 14A is free-form rather than SEC-item-addressable. Fetch the full filing, inspect its +opening table of contents, then grep its Markdown for the filing's own headings: + +```bash +python scripts/fetch_filing.py --ticker AAPL --form "DEF 14A" +``` + +Prefer an exact accession when orientation reports several proxy-related filings. Common +heading language varies, so search related terms rather than assuming one exact title. + +## Beneficial ownership + +Look for headings such as: + +- Security Ownership of Certain Beneficial Owners and Management; +- Principal Shareholders; +- Beneficial Ownership; +- Stock Ownership of Directors and Executive Officers. + +Capture the measurement date, share class, shares, percentage, footnotes, and whether the row +covers a person, reporting group, or aggregate management. Do not combine 13F holdings with +proxy ownership percentages without reconciling dates, classes, and reporting entities. + +For external 5% schedules and institutional filings, use `guide_holdings.md`; the proxy table +is the issuer's annual consolidated disclosure and may have a different measurement date. + +## Board and committees + +Relevant sections commonly cover: + +- director biographies, tenure, qualifications, and other directorships; +- independence determinations; +- audit, compensation, and nominating/governance committee membership; +- board leadership and lead-independent-director structure; +- meeting attendance; +- risk oversight; +- director compensation. + +Extract what the filing states and preserve the relevant date. Avoid turning structural facts +into a governance score inside this data skill. + +## Executive compensation + +Regulation S-K Item 402 disclosures commonly include: + +- Compensation Discussion and Analysis; +- Summary Compensation Table; +- Grants of Plan-Based Awards; +- Outstanding Equity Awards at Fiscal Year-End; +- Option Exercises and Stock Vested; +- Pension Benefits and Nonqualified Deferred Compensation; +- Potential Payments upon Termination or Change in Control; +- Pay Versus Performance; +- CEO pay ratio. + +Capture table periods, units, footnotes, grant terms, and award status. “Total compensation” is +not interchangeable with realized pay or current equity value. + +## Related parties, auditors, and equity plans + +Search for: + +- Certain Relationships and Related Transactions; +- Transactions with Related Persons; +- independent registered public accounting firm; +- audit and non-audit fees; +- auditor ratification; +- Equity Compensation Plan Information; +- request to approve or amend an equity incentive plan. + +For equity plans, preserve authorized, outstanding, and remaining-share definitions rather +than collapsing them into one dilution number. Securities offerings, warrants, convertibles, +and ATM programs may instead require 10-Q/10-K footnotes, 8-K exhibits, S-3, or 424B filings. + +## Proposals and voting + +The proxy explains each management or shareholder proposal and the board's stated +recommendation. The eventual vote is generally reported in Form 8-K Item 5.07: + +```bash +python scripts/fetch_filing.py --ticker AAPL --form 8-K --section "Item 5.07" +``` + +Keep proposals, recommendations, and final voting results distinct. Record votes for, against, +abstained, and broker non-votes as presented rather than reducing them to a pass/fail label. + +## Foreign private issuers + +Foreign private issuers generally do not file DEF 14A. Start with Form 20-F Item 6: + +- Item 6.A — directors and senior management; +- Item 6.B — compensation; +- Item 6.C — board practices; +- Item 6.D — employees; +- Item 6.E — share ownership. + +Home-country annual reports, meeting circulars, or regulator filings may contain greater +detail. Form 20-F Item 7.A covers major shareholders; when slicing by text, terminate at Item 8 +rather than the words “Related Party Transactions,” which also occur in Item 7's full title. diff --git a/skills/sec-edgar-skill/scripts/_common.py b/skills/sec-edgar-skill/scripts/_common.py index 877e218..3c64656 100644 --- a/skills/sec-edgar-skill/scripts/_common.py +++ b/skills/sec-edgar-skill/scripts/_common.py @@ -1,164 +1,207 @@ -"""Shared bootstrap and conventions for the sec-edgar-skill scripts. - -Importing this module gives every script identical, correct runtime setup with -zero duplication: - * UTF-8 stdout/stderr on Windows (edgartools' rich reprs contain emoji that - raise UnicodeEncodeError on cp1252 consoles), - * truststore injected into SSL when available (so HTTPS works behind an - inspecting corporate proxy instead of raising CERTIFICATE_VERIFY_FAILED). - -It also centralises the two contracts that must stay identical across every -script: the SEC identity requirement, and the on-disk cache layout. Keeping -them here is why a filing cached by one script is found by the others. - -Output convention: human-readable progress goes to stderr via ``log``; the -machine-readable result (always an absolute path) goes to stdout via ``emit``, -so a calling agent can capture the path without parsing log noise. -""" - -from __future__ import annotations - -import argparse -import os -import sys -from pathlib import Path - -# --- Runtime setup (runs once, on import) --------------------------------- -if sys.platform.startswith("win"): - for _stream in (sys.stdout, sys.stderr): - try: - _stream.reconfigure(encoding="utf-8") - except Exception: - pass - -try: - import truststore - - truststore.inject_into_ssl() -except Exception: - pass - - -def log(msg: str) -> None: - """Human-readable progress -> stderr (keeps stdout clean for results).""" - print(msg, file=sys.stderr, flush=True) - - -def emit(path: str | os.PathLike) -> None: - """Machine-readable result -> stdout: one absolute path per line.""" - print(str(Path(path).resolve())) - - -# --- SEC identity: a mechanical requirement, never a silent default ------- -def resolve_identity(cli_value: str | None = None) -> str: - """Return the SEC identity or exit(2) with an actionable message. - - The SEC fair-access policy requires a real ``Name email`` User-Agent; - requests without one are blocked with HTTP 403. We deliberately do NOT - fall back to a fake default, because sending bogus contact information - abuses the policy and risks the host IP being rate-limited or blocked. - """ - identity = cli_value or os.environ.get("EDGAR_IDENTITY") - if not identity or "@" not in identity: - log( - "ERROR: a SEC identity is required (SEC fair-access policy; missing " - "it returns HTTP 403). Set it once:\n" - ' PowerShell: $env:EDGAR_IDENTITY = "Jane Analyst jane@example.com"\n' - ' Bash: export EDGAR_IDENTITY="Jane Analyst jane@example.com"\n' - 'or pass --identity "Name email@example.com".' - ) - sys.exit(2) - from edgar import set_identity - - set_identity(identity) - return identity - - -def add_identity_arg(parser: argparse.ArgumentParser) -> None: - parser.add_argument( - "--identity", - help='SEC User-Agent "Name email@example.com" (else uses $EDGAR_IDENTITY).', - ) - - -# --- Cache contract ------------------------------------------------------- -def cache_root(cli_value: str | None = None) -> Path: - """Resolve the cache root: --cache-dir > $SEC_CACHE_DIR > ./sec-cache. - - Default is workspace-relative and *visible* on purpose: - * inside the working directory, so the agent's native grep/ripgrep finds - cached files with no absolute-path gymnastics; - * persistent across runs, unlike an OS temp dir (which gets purged, which - would defeat the whole point of caching and re-hit the SEC needlessly); - * not dot-prefixed, because ripgrep skips dotdirs by default. - Set $SEC_CACHE_DIR to redirect to an OS cache dir if you want no footprint. - """ - root = cli_value or os.environ.get("SEC_CACHE_DIR") or "sec-cache" - return Path(root).resolve() - - -def add_cache_arg(parser: argparse.ArgumentParser) -> None: - parser.add_argument( - "--cache-dir", - help="Cache root (default: $SEC_CACHE_DIR or ./sec-cache).", - ) - - -def safe_component(part: object) -> str: - """Make a single path component filesystem-safe (e.g. 10-K/A -> 10-K-A).""" - cleaned = "".join(ch if (ch.isalnum() or ch in "-._") else "-" for ch in str(part)) - return cleaned.strip("-") or "x" - - -def filing_stem(filing) -> str: - """Deterministic, collision-free stem: ``FORM_FILINGDATE_ACCESSION``. - - Accession numbers are globally unique and verifiable against SEC, so this - lets a caller answer "is this already cached?" by globbing the company - directory instead of re-downloading, and avoids the stale ``latest_*.md`` - anti-pattern (a "latest" name silently goes out of date next quarter). - """ - parts = [ - safe_component(getattr(filing, "form", "filing")), - safe_component(getattr(filing, "filing_date", "")), - safe_component( - getattr(filing, "accession_no", None) or getattr(filing, "accession_number", "") - ), - ] - return "_".join(p for p in parts if p and p != "x") - - -def company_dir(root: Path, company, ticker_hint: str | None = None) -> Path: - """Return (and create) ``//``; falls back to CIK if unknown.""" - label = ticker_hint - if not label: - try: - tickers = getattr(company, "tickers", None) - if tickers: - label = next(iter(tickers)) - except Exception: - label = None - if not label: - label = str(getattr(company, "cik", "unknown")) - directory = root / safe_component(label).upper() - directory.mkdir(parents=True, exist_ok=True) - return directory - - -def resolve_company(ticker_or_cik: str): - """Look up a Company, exiting(1) with a clean message if it can't resolve.""" - from edgar import Company - - try: - return Company(ticker_or_cik) - except Exception as exc: # surface one clean line to the calling agent - log(f"ERROR: could not resolve company '{ticker_or_cik}': {exc}") - sys.exit(1) - - -def write_text(path: str | os.PathLike, content: str) -> Path: - """Write UTF-8 text, creating parent directories as needed.""" - p = Path(path) - p.parent.mkdir(parents=True, exist_ok=True) - p.write_text(content, encoding="utf-8") - return p +"""Shared bootstrap and conventions for the sec-edgar-skill scripts. + +Importing this module gives every script identical, correct runtime setup with +zero duplication: + * UTF-8 stdout/stderr on Windows (edgartools' rich reprs contain emoji that + raise UnicodeEncodeError on cp1252 consoles), + * truststore injected into SSL when available (so HTTPS works behind an + inspecting corporate proxy instead of raising CERTIFICATE_VERIFY_FAILED). + +It also centralises the two contracts that must stay identical across every +script: the SEC identity requirement, and the on-disk cache layout. Keeping +them here is why a filing cached by one script is found by the others. + +Output convention: human-readable progress goes to stderr via ``log``. Commands +that create an artifact send one absolute path per line to stdout via ``emit``; +inspection commands may instead print their report to stdout. +""" + +from __future__ import annotations + +import argparse +import os +import sys +from pathlib import Path + +# --- Runtime setup (runs once, on import) --------------------------------- +if sys.platform.startswith("win"): + for _stream in (sys.stdout, sys.stderr): + try: + _stream.reconfigure(encoding="utf-8") + except Exception: + pass + +try: + import truststore + + truststore.inject_into_ssl() +except Exception: + pass + + +def log(msg: str) -> None: + """Human-readable progress -> stderr (keeps stdout clean for results).""" + print(msg, file=sys.stderr, flush=True) + + +def emit(path: str | os.PathLike) -> None: + """Machine-readable result -> stdout: one absolute path per line.""" + print(str(Path(path).resolve())) + + +# --- SEC identity: a mechanical requirement, never a silent default ------- +def resolve_identity(cli_value: str | None = None) -> str: + """Return the SEC identity or exit(2) with an actionable message. + + The SEC fair-access policy requires a real ``Name email`` User-Agent; + requests without one are blocked with HTTP 403. We deliberately do NOT + fall back to a fake default, because sending bogus contact information + abuses the policy and risks the host IP being rate-limited or blocked. + """ + identity = cli_value or os.environ.get("EDGAR_IDENTITY") + if not identity or "@" not in identity: + log( + "ERROR: a SEC identity is required (SEC fair-access policy; missing " + "it returns HTTP 403). Set it once:\n" + ' PowerShell: $env:EDGAR_IDENTITY = "Jane Analyst jane@example.com"\n' + ' Bash: export EDGAR_IDENTITY="Jane Analyst jane@example.com"\n' + 'or pass --identity "Name email@example.com".' + ) + sys.exit(2) + from edgar import set_identity + + set_identity(identity) + return identity + + +def add_identity_arg(parser: argparse.ArgumentParser) -> None: + parser.add_argument( + "--identity", + help='SEC User-Agent "Name email@example.com" (else uses $EDGAR_IDENTITY).', + ) + + +# --- Cache contract ------------------------------------------------------- +def cache_root(cli_value: str | None = None) -> Path: + """Resolve the cache root: --cache-dir > $SEC_CACHE_DIR > ./sec-cache. + + Default is workspace-relative and *visible* on purpose: + * inside the working directory, so the agent's native grep/ripgrep finds + cached files with no absolute-path gymnastics; + * persistent across runs, unlike an OS temp dir (which gets purged, which + would defeat the whole point of caching and re-hit the SEC needlessly); + * not dot-prefixed, because ripgrep skips dotdirs by default. + Set $SEC_CACHE_DIR to redirect to an OS cache dir if you want no footprint. + """ + root = cli_value or os.environ.get("SEC_CACHE_DIR") or "sec-cache" + return Path(root).resolve() + + +def add_cache_arg(parser: argparse.ArgumentParser) -> None: + parser.add_argument( + "--cache-dir", + help="Cache root (default: $SEC_CACHE_DIR or ./sec-cache).", + ) + + +def add_force_arg(parser: argparse.ArgumentParser) -> None: + """Add the common override for immutable, accession-keyed artifacts.""" + parser.add_argument( + "--force", + action="store_true", + help="Regenerate an existing accession-keyed artifact.", + ) + + +def emit_cached(path: str | os.PathLike, *, force: bool = False) -> bool: + """Emit a non-empty cached artifact unless regeneration was requested.""" + cached = Path(path) + if not force and cached.is_file() and cached.stat().st_size > 0: + log(f"Using cached artifact: {cached.resolve()}") + emit(cached) + return True + return False + + +def safe_component(part: object) -> str: + """Make a single path component filesystem-safe (e.g. 10-K/A -> 10-K-A).""" + cleaned = "".join(ch if (ch.isalnum() or ch in "-._") else "-" for ch in str(part)) + return cleaned.strip("-") or "x" + + +def filing_stem(filing) -> str: + """Deterministic, collision-free stem: ``FORM_FILINGDATE_ACCESSION``. + + Accession numbers are globally unique and verifiable against SEC, so this + lets a caller answer "is this already cached?" by globbing the company + directory instead of re-downloading, and avoids the stale ``latest_*.md`` + anti-pattern (a "latest" name silently goes out of date next quarter). + """ + parts = [ + safe_component(getattr(filing, "form", "filing")), + safe_component(getattr(filing, "filing_date", "")), + safe_component( + getattr(filing, "accession_no", None) or getattr(filing, "accession_number", "") + ), + ] + return "_".join(p for p in parts if p and p != "x") + + +def company_dir(root: Path, company, ticker_hint: str | None = None) -> Path: + """Return (and create) ``//``; falls back to CIK if unknown.""" + label = ticker_hint + if not label: + try: + tickers = getattr(company, "tickers", None) + if tickers: + label = next(iter(tickers)) + except Exception: + label = None + if not label: + label = str(getattr(company, "cik", "unknown")) + directory = root / safe_component(label).upper() + directory.mkdir(parents=True, exist_ok=True) + return directory + + +def resolve_company(ticker_or_cik: str): + """Look up a Company, exiting(1) with a clean message if it can't resolve.""" + from edgar import Company + + try: + return Company(ticker_or_cik) + except Exception as exc: # surface one clean line to the calling agent + log(f"ERROR: could not resolve company '{ticker_or_cik}': {exc}") + sys.exit(1) + + +def resolve_filing(accession: str): + """Resolve one filing by accession number, exiting cleanly on failure.""" + from edgar import find + + try: + filing = find(accession) + except Exception as exc: + log(f"ERROR: could not resolve accession '{accession}': {exc}") + sys.exit(1) + if filing is None or not hasattr(filing, "accession_no"): + log(f"ERROR: accession '{accession}' did not resolve to an SEC filing.") + sys.exit(1) + return filing + + +def company_for_filing(filing): + """Resolve the filing's registrant so cache paths retain a ticker when possible.""" + cik = getattr(filing, "cik", None) + if cik is None: + log("ERROR: the resolved filing does not identify a registrant CIK.") + sys.exit(1) + return resolve_company(str(cik)) + + +def write_text(path: str | os.PathLike, content: str) -> Path: + """Write UTF-8 text, creating parent directories as needed.""" + p = Path(path) + p.parent.mkdir(parents=True, exist_ok=True) + p.write_text(content, encoding="utf-8") + return p diff --git a/skills/sec-edgar-skill/scripts/fetch_13f_holders.py b/skills/sec-edgar-skill/scripts/fetch_13f_holders.py index eba9e34..777e30e 100644 --- a/skills/sec-edgar-skill/scripts/fetch_13f_holders.py +++ b/skills/sec-edgar-skill/scripts/fetch_13f_holders.py @@ -1,39 +1,25 @@ -"""Fetch institutional 13F holder data from 13f.info for a stock or manager. +"""Fetch distilled SEC Form 13F data for a stock, manager, or manager-stock pair. -13f.info is a free, structured interface to SEC 13F filings that provides -pre-parsed, queryable data far more efficiently than parsing raw 13F XML from -EDGAR. It offers three key views: - - **Stock-centric** (who owns this stock?): - --ticker AAPL → all institutional holders for the latest quarter - --ticker AAPL --history → quarterly holder/share-count summary over time - - **Manager-centric** (what does this fund hold?): - --manager "Berkshire Hathaway" → search for a manager, show their holdings - --cik 0001067983 → look up a manager by CIK directly - - **Cross-reference** (how has a specific manager's position changed?): - --cik 0000906304 --cusip 205826209 → one manager's history with one stock - -All output is written to the cache and the absolute path is emitted to stdout: - - Stock-centric: ``//13f-holders_-Q.md`` - (or ``13f-history_.md`` for --history) - - Manager-centric: ``/managers/13f-manager_.md`` - - Cross-reference: ``//13f-xref_.md`` - (or ``/managers/13f-xref__.md`` without --ticker) - -Data source: https://13f.info (public, no API key needed). +The normal interface accepts tickers and manager names (or exact manager CIKs), +resolving lower-level CUSIPs internally. Reports expose the underlying SEC period, +CIK, and accession rather than the retrieval provider. No SEC identity is required. """ from __future__ import annotations import argparse import json +import re import sys +from dataclasses import dataclass +from datetime import date, datetime, timedelta +from html import unescape from urllib.error import HTTPError, URLError +from urllib.parse import quote from urllib.request import Request, urlopen import _common as c +import pandas as pd _BASE = "https://13f.info" _UA = ( @@ -42,519 +28,529 @@ ) -# --------------------------------------------------------------------------- -# HTTP / JSON helpers -# --------------------------------------------------------------------------- +class FetchFailure(RuntimeError): + """Raised when the distilled 13F backend cannot complete a request.""" -def _get_json(url: str) -> dict | list | None: - """Fetch a JSON endpoint, return parsed data or None on failure.""" - req = Request(url, headers={"User-Agent": _UA, "Accept": "application/json"}) - try: - with urlopen(req, timeout=30) as resp: - return json.loads(resp.read().decode("utf-8", errors="replace")) - except (HTTPError, URLError, json.JSONDecodeError) as exc: - c.log(f"WARNING: fetch failed for {url}: {exc}") - return None +@dataclass +class ManagerFiling: + """One filing row and its linked distilled portfolio page.""" - -def _get_html(url: str) -> str: - """Fetch an HTML page and return the body as a string.""" - req = Request(url, headers={"User-Agent": _UA, "Accept": "text/html"}) - try: - with urlopen(req, timeout=30) as resp: - return resp.read().decode("utf-8", errors="replace") - except (HTTPError, URLError) as exc: - c.log(f"WARNING: fetch failed for {url}: {exc}") - return "" + url: str + period: str + holdings: str + form_type: str + filed: str + filing_id: str -# --------------------------------------------------------------------------- -# API endpoints (discovered from 13f.info's client-side code) -# --------------------------------------------------------------------------- -# Working JSON endpoints: -# /data/autocomplete?q=... → search managers + CUSIPs -# /data/cusip/{cusip}/{year}/{quarter} → all holders for a CUSIP in a quarter -# /data/manager/{cik}/cusip/{cusip} → one manager's history with one CUSIP -# -# HTML-only (no JSON API): -# /cusip/{cusip} → stock overview (quarterly summary table) -# /manager/{cik}-slug → manager overview (filing history table) -# /13f/{filing-slug} → single filing holdings table - +def _get_json(url: str) -> dict | list: + request = Request(url, headers={"User-Agent": _UA, "Accept": "application/json"}) + try: + with urlopen(request, timeout=30) as response: + return json.loads(response.read().decode("utf-8", errors="replace")) + except (HTTPError, URLError, TimeoutError, json.JSONDecodeError) as exc: + raise FetchFailure(f"13F data request failed: {exc}") from exc -def _autocomplete(query: str) -> dict | None: - """Search for managers and CUSIPs by name/ticker.""" - from urllib.parse import quote - return _get_json(f"{_BASE}/data/autocomplete?q={quote(query)}") +def _get_html(url: str) -> str: + request = Request(url, headers={"User-Agent": _UA, "Accept": "text/html"}) + try: + with urlopen(request, timeout=30) as response: + return response.read().decode("utf-8", errors="replace") + except (HTTPError, URLError, TimeoutError) as exc: + raise FetchFailure(f"13F page request failed: {exc}") from exc -def _holders_for_quarter(cusip: str, year: int, quarter: int) -> dict | None: - """Get all institutional holders for a CUSIP in a specific quarter.""" - return _get_json(f"{_BASE}/data/cusip/{cusip}/{year}/{quarter}") +def _autocomplete(query: str) -> dict: + data = _get_json(f"{_BASE}/data/autocomplete?q={quote(query)}") + return data if isinstance(data, dict) else {} -def _manager_cusip_history(cik: str, cusip: str) -> dict | None: - """Get a specific manager's position history for a specific CUSIP.""" - return _get_json(f"{_BASE}/data/manager/{cik}/cusip/{cusip}") +def _holders_for_quarter(cusip: str, year: int, quarter: int) -> dict: + data = _get_json(f"{_BASE}/data/cusip/{cusip}/{year}/{quarter}") + return data if isinstance(data, dict) else {} -# --------------------------------------------------------------------------- -# Resolve ticker → CUSIP via autocomplete -# --------------------------------------------------------------------------- +def _manager_cusip_history(cik: str, cusip: str) -> dict: + data = _get_json(f"{_BASE}/data/manager/{cik}/cusip/{cusip}") + return data if isinstance(data, dict) else {} def _resolve_cusip(ticker: str) -> tuple[str, str, str] | None: - """Resolve a ticker to (cusip, symbol, issuer_name) via autocomplete. - - Returns the first equity CUSIP match, or None. - """ data = _autocomplete(ticker) - if not data or not data.get("cusips"): - return None - for entry in data["cusips"]: + for entry in data.get("cusips", []): name = entry.get("name", "") extra = entry.get("extra", "") - # Format: "AAPL - Apple Inc." / "037833100 - COM" parts = name.split(" - ", 1) - sym = parts[0].strip() if parts else "" + symbol = parts[0].strip() if parts else "" issuer = parts[1].strip() if len(parts) > 1 else name cusip = extra.split(" - ")[0].strip() if " - " in extra else extra.strip() - if cusip and sym.upper() == ticker.upper(): - return (cusip, sym, issuer) - # Fallback: return the first CUSIP result - entry = data["cusips"][0] - name = entry.get("name", "") - extra = entry.get("extra", "") - parts = name.split(" - ", 1) - sym = parts[0].strip() if parts else "" - issuer = parts[1].strip() if len(parts) > 1 else name - cusip = extra.split(" - ")[0].strip() if " - " in extra else extra.strip() - return (cusip, sym, issuer) - - -def _resolve_manager(query: str) -> tuple[str, str] | None: - """Resolve a manager name to (cik, name) via autocomplete. - - Returns the first manager match, or None. - """ + if cusip and symbol.upper() == ticker.upper(): + return cusip, symbol, issuer + return None + + +def _manager_name_from_html(html: str, fallback: str) -> str: + match = re.search(r"]*>(.*?)", html, re.DOTALL | re.IGNORECASE) + if not match: + return fallback + return re.sub(r"<[^>]+>", "", unescape(match.group(1))).strip() or fallback + + +def _resolve_manager(query: str) -> tuple[str, str, str, str] | None: + digits = "".join(character for character in query if character.isdigit()) + if len(digits) == 10 and query.strip().replace("-", "").isdigit(): + url = f"/manager/{digits}" + html = _get_html(f"{_BASE}{url}") + return digits, _manager_name_from_html(html, digits), url, html + data = _autocomplete(query) - if not data or not data.get("managers"): + managers = data.get("managers", []) + if not managers: return None - entry = data["managers"][0] - name = entry.get("name", "") - url = entry.get("url", "") - # URL format: /manager/0001067983-berkshire-hathaway-inc + + normalized = query.casefold().strip() + exact = [ + entry for entry in managers if str(entry.get("name", "")).casefold().strip() == normalized + ] + candidates = [ + entry for entry in managers if normalized in str(entry.get("name", "")).casefold() + ] or managers + if len(exact) == 1: + entry = exact[0] + elif len(candidates) == 1: + entry = candidates[0] + else: + choices = [] + for candidate in candidates[:5]: + candidate_url = str(candidate.get("url", "")) + candidate_cik = candidate_url.split("/")[-1].split("-")[0] + choices.append(f"{candidate.get('name', '')} ({candidate_cik})") + raise FetchFailure( + f"manager name '{query}' is ambiguous; rerun --manager with one of these CIKs: " + + "; ".join(choices) + ) + name = str(entry.get("name", "")).strip() + url = str(entry.get("url", "")).strip() cik = url.split("/")[-1].split("-")[0] if url else "" - return (cik, name) + if not cik or not url: + return None + html = _get_html(f"{_BASE}{url}") + return cik, name or _manager_name_from_html(html, query), url, html -# --------------------------------------------------------------------------- -# Formatting helpers -# --------------------------------------------------------------------------- +def _fmt_shares(value: object) -> str: + try: + if pd.isna(value): + return "n/a" + number = float(str(value).replace(",", "")) + except (TypeError, ValueError): + return str(value) + if abs(number) >= 1_000_000: + return f"{number / 1_000_000:,.2f}M" + if abs(number) >= 1_000: + return f"{number / 1_000:,.1f}K" + return f"{number:,.0f}" + + +def _fmt_percent(value: object) -> str: + try: + if pd.isna(value): + return "n/a" + return f"{float(value):,.1f}%" + except (TypeError, ValueError): + return "n/a" -def _fmt_shares(s: int | float | None) -> str: - if s is None: - return "n/a" - if s >= 1_000_000: - return f"{s / 1_000_000:,.2f}M" - if s >= 1_000: - return f"{s / 1_000:,.1f}K" - return f"{s:,}" +def _cell(value: object) -> str: + try: + if pd.isna(value): + return "" + except (TypeError, ValueError): + pass + return str(value).replace("|", "\\|").replace("\n", " ").strip() -def _fmt_pct(p: float | None) -> str: - if p is None: - return "n/a" - return f"{p:.1f}%" +def _accession_digits(value: object) -> str: + match = re.search(r"(?:/13f/)?(\d{18})", str(value or "")) + return match.group(1) if match else "" + + +def _format_accession(value: object) -> str: + digits = _accession_digits(value) + if len(digits) == 18: + return f"{digits[:10]}-{digits[10:12]}-{digits[12:]}" + return str(value or "") + + +def _iso_date(value: str) -> str: + try: + return datetime.strptime(value, "%m/%d/%Y").date().isoformat() + except ValueError: + return value -# --------------------------------------------------------------------------- -# Output: Stock-centric (who owns this stock?) -# --------------------------------------------------------------------------- +def _quarter_end(year: int, quarter: int) -> str: + return { + 1: f"{year}-03-31", + 2: f"{year}-06-30", + 3: f"{year}-09-30", + 4: f"{year}-12-31", + }[quarter] + + +def _latest_quarter() -> tuple[int, int]: + today = date.today() + quarters = [] + for year in (today.year, today.year - 1): + quarters.extend( + [ + (date(year, 3, 31), year, 1), + (date(year, 6, 30), year, 2), + (date(year, 9, 30), year, 3), + (date(year, 12, 31), year, 4), + ] + ) + for end_date, year, quarter in sorted(quarters, reverse=True): + if today - end_date >= timedelta(days=50): + return year, quarter + return today.year - 1, 4 def _build_stock_holders( ticker: str, cusip: str, issuer: str, year: int, quarter: int, top_n: int ) -> str | None: - """Build Markdown for the top institutional holders in a given quarter. - - Returns the Markdown string, or None if no data is available. - """ data = _holders_for_quarter(cusip, year, quarter) - if not data or not data.get("data"): - c.log(f"No holder data found for {ticker} ({cusip}) Q{quarter} {year}.") + holders = data.get("data", []) + if not holders: return None - holders = data["data"] - lines = [] - - # Determine period end date from the first entry - period_end = "" - if holders and isinstance(holders[0][1], list): - period_end = holders[0][1][0] # e.g. "2026-03-31" - - lines.append(f"# 13F Institutional Holders: {ticker} ({issuer})") - lines.append("") - lines.append(f"- **CUSIP:** {cusip}") - lines.append(f"- **Period:** Q{quarter} {year}") - if period_end: - lines.append(f"- **Holdings as of:** {period_end}") - lines.append(f"- **Total holders reporting:** {len(holders)}") - lines.append(f"- **Source:** {_BASE}/cusip/{cusip}/{year}/{quarter}") - lines.append("") - - # Each holder entry: [[manager_name, cik, cusip], [date, filing_slug], value, shares, principal] - lines.append(f"## Top {min(top_n, len(holders))} Holders by Shares") - lines.append("") - lines.append("| # | Manager | Shares | CIK |") - lines.append("|---|---------|--------|-----|") - - # Sort by shares descending - ranked = sorted(holders, key=lambda h: h[3] or 0, reverse=True) - - for i, h in enumerate(ranked[:top_n]): - manager_info = h[0] - shares = h[3] - + period_end = _quarter_end(year, quarter) + if isinstance(holders[0][1], list) and holders[0][1]: + period_end = holders[0][1][0] + ranked = sorted(holders, key=lambda holder: holder[3] or 0, reverse=True) + lines = [ + f"# 13F Institutional Holders: {ticker} ({issuer})", + "", + f"- **CUSIP:** {cusip}", + f"- **Reporting period:** Q{quarter} {year}", + f"- **Holdings as of:** {period_end}", + f"- **Reporting managers found:** {len(holders)}", + "- **Underlying records:** SEC Form 13F filings", + "", + f"## Top {min(top_n, len(holders))} Reporting Managers by Shares", + "", + "| # | Manager | Shares | CIK | SEC Accession |", + "|---:|---|---:|---|---|", + ] + for index, holder in enumerate(ranked[:top_n], 1): + manager_info = holder[0] + filing_info = holder[1] name = manager_info[0] if isinstance(manager_info, list) else str(manager_info) cik = manager_info[1] if isinstance(manager_info, list) and len(manager_info) > 1 else "" - - lines.append(f"| {i + 1} | {name} | {_fmt_shares(shares)} | {cik} |") - + filing_slug = ( + filing_info[1] if isinstance(filing_info, list) and len(filing_info) > 1 else "" + ) + lines.append( + f"| {index} | {_cell(name)} | {_fmt_shares(holder[3])} | {cik} " + f"| {_format_accession(filing_slug)} |" + ) if len(holders) > top_n: - lines.append("") - lines.append(f"*Showing top {top_n} of {len(holders)} holders. Use --top to adjust.*") - + lines.extend(["", f"*Showing {top_n} of {len(holders)} reporting managers.*"]) return "\n".join(lines) def _build_stock_history(ticker: str, cusip: str, issuer: str) -> str | None: - """Build Markdown for the quarterly holder/share-count history. - - Returns the Markdown string, or None if no data is available. - This requires scraping the HTML page since there's no JSON endpoint - for the CUSIP overview. - """ - import re - html = _get_html(f"{_BASE}/cusip/{cusip}") - if not html: - c.log(f"No history data found for {ticker} ({cusip}).") - return None - - lines = [] - lines.append(f"# 13F Holder History: {ticker} ({issuer})") - lines.append("") - lines.append(f"- **CUSIP:** {cusip}") - lines.append(f"- **Source:** {_BASE}/cusip/{cusip}") - lines.append("") - - # Parse the HTML table rows — each row has 5 cells: - # period (with link), filings, shares, value, options value - row_pattern = re.compile( + pattern = re.compile( r"]*>\s*" - r"]*>\s*]*>(\d{4}\s+Q\d)\s*\s*" # period - r"]*>\s*(\d+)\s*\s*" # filings - r"]*>\s*([^<]+?)\s*\s*" # shares - r"]*>\s*([^<]+?)\s*", # value + r"]*>\s*]*>(\d{4}\s+Q\d)\s*\s*" + r"]*>\s*([\d,]+)\s*\s*" + r"]*>\s*([^<]+?)\s*", re.DOTALL, ) - - lines.append("| Period | Holders | Shares (excl. options) |") - lines.append("|--------|---------|----------------------|") - - for m in row_pattern.finditer(html): - period = m.group(1).strip() - filings = m.group(2).strip() - shares = m.group(3).strip() - lines.append(f"| {period} | {filings} | {shares} |") - + rows = [tuple(value.strip() for value in match.groups()) for match in pattern.finditer(html)] + if not rows: + return None + lines = [ + f"# 13F Holder History: {ticker} ({issuer})", + "", + f"- **CUSIP:** {cusip}", + "- **Underlying records:** SEC Form 13F filings", + "", + "| Period | Reporting Managers | Shares (excluding options) |", + "|---|---:|---:|", + ] + lines.extend(f"| {period} | {filings} | {shares} |" for period, filings, shares in rows) return "\n".join(lines) -# --------------------------------------------------------------------------- -# Output: Manager-centric (what does this fund hold?) -# --------------------------------------------------------------------------- - - -def _build_manager_holdings(cik: str, manager_name: str) -> str | None: - """Build Markdown for a manager's filing history (scraped from HTML). +def _parse_manager_filings(html: str) -> list[ManagerFiling]: + pattern = re.compile( + r']*href="(/13f/[^"]+)"[^>]*>\s*(Q\d\s+\d{4})\s*' + r".*?]*>\s*(\d+)\s*" + r".*?]*>\s*([\d,]+)\s*" + r'.*?]*title="([^"]*)"[^>]*>.*?' + r'.*?]*title="([^"]*)"[^>]*>.*?' + r".*?]*>\s*([\d/]+)\s*", + re.DOTALL, + ) + filings = [] + for match in pattern.finditer(html): + url, period, holdings, _value, _top, form_type, filed = match.groups() + filings.append( + ManagerFiling( + url=url, + period=period, + holdings=holdings, + form_type=form_type, + filed=filed, + filing_id=url, + ) + ) + return filings - Returns the Markdown string, or None if no data is available. - """ - import re - # Try to find the manager's page URL via autocomplete - data = _autocomplete(manager_name) - manager_url = None - if data and data.get("managers"): - for m in data["managers"]: - url = m.get("url", "") - if cik in url: - manager_url = url - break - if not manager_url: - manager_url = data["managers"][0].get("url", "") +def _select_manager_filing( + filings: list[ManagerFiling], year: int | None, quarter: int | None +) -> ManagerFiling | None: + eligible = [ + filing for filing in filings if filing.form_type.upper() in {"13F-HR", "RESTATEMENT"} + ] + if year is not None and quarter is not None: + target = f"Q{quarter} {year}" + eligible = [filing for filing in eligible if filing.period == target] + return eligible[0] if eligible else None - if not manager_url: - manager_url = f"/manager/{cik}" - html = _get_html(f"{_BASE}{manager_url}") - if not html: - c.log(f"No data found for manager {manager_name} ({cik}).") +def _build_manager_holdings(cik: str, manager_name: str, filing: ManagerFiling) -> str | None: + filing_digits = _accession_digits(filing.url) + if not filing_digits: + return None + data = _get_json(f"{_BASE}/data/13f/{filing_digits}") + rows = data.get("data", []) if isinstance(data, dict) else [] + if not rows: return None - lines = [] - lines.append(f"# 13F Filings: {manager_name}") - lines.append("") - lines.append(f"- **CIK:** {cik}") - lines.append(f"- **Source:** {_BASE}{manager_url}") - lines.append("") + lines = [ + f"# 13F Portfolio: {manager_name}", + "", + f"- **Manager CIK:** {cik}", + f"- **Reporting period:** {filing.period}", + f"- **Filed:** {_iso_date(filing.filed)}", + f"- **SEC accession:** {_format_accession(filing.filing_id)}", + "- **Underlying record:** SEC Form 13F", + "", + f"## Disclosed Holdings ({len(rows)})", + "", + "| Symbol | Issuer | Class | CUSIP | Shares/Principal | Option Type |", + "|---|---|---|---|---:|---|", + ] + # Row shape: symbol, issuer, class, CUSIP, value, portfolio %, shares, + # principal, option type. Dollar value and portfolio percentage are omitted. + for row in sorted(rows, key=lambda item: str(item[0] or "")): + symbol, issuer, security_class, cusip = row[:4] + shares = row[6] if len(row) > 6 else None + principal = row[7] if len(row) > 7 else None + option_type = row[8] if len(row) > 8 else None + amount = shares if shares is not None else principal + lines.append( + f"| {_cell(symbol)} | {_cell(issuer)} | {_cell(security_class)} | {_cell(cusip)} " + f"| {_fmt_shares(amount)} | {_cell(option_type)} |" + ) + return "\n".join(lines) - # Parse the filing history table — 7 columns: - # Quarter (link), Holdings, Value, Top Holdings (title attr), Form Type, Date Filed, Filing ID - row_pattern = re.compile( - r']*href="(/13f/[^"]+)"[^>]*>\s*(Q\d\s+\d{4})\s*' - r".*?]*>\s*(\d+)\s*" # holdings - r".*?]*>\s*([\d,]+)\s*" # value - r'.*?]*title="([^"]*)"[^>]*>.*?' # top holdings (title attr) - r'.*?]*title="([^"]*)"[^>]*>.*?' # form type (title attr) - r".*?]*>\s*([\d/]+)\s*", # date filed - re.DOTALL, - ) - lines.append("| Quarter | Holdings | Top Holdings | Type | Filed |") - lines.append("|---------|----------|--------------|------|-------|") - - count = 0 - for m in row_pattern.finditer(html): - quarter = m.group(2).strip() - holdings = m.group(3).strip() - top = m.group(5).strip() - form_type = m.group(6).strip() - filed = m.group(7).strip() - lines.append(f"| {quarter} | {holdings} | {top} | {form_type} | {filed} |") - count += 1 - if count >= 20: - break - - if count == 0: - c.log("WARNING: could not parse manager filing history from HTML.") +def _build_manager_history(cik: str, manager_name: str, filings: list[ManagerFiling]) -> str | None: + complete = [ + filing for filing in filings if filing.form_type.upper() in {"13F-HR", "RESTATEMENT"} + ] + if not complete: return None + lines = [ + f"# 13F Portfolio History: {manager_name}", + "", + f"- **Manager CIK:** {cik}", + "- **Underlying records:** SEC Form 13F filings", + ] + if len(complete) >= 2: + current, previous = complete[:2] + current_id = _accession_digits(current.filing_id) + previous_id = _accession_digits(previous.filing_id) + comparison = _get_json(f"{_BASE}/data/13f/{current_id}/compare/{previous_id}") + rows = comparison.get("data", []) if isinstance(comparison, dict) else [] + if rows: + lines.extend( + [ + "", + f"## Share Changes: {previous.period} to {current.period}", + "", + "| Symbol | Issuer | Class | CUSIP | Option | Previous | Current | Change | Change % |", + "|---|---|---|---|---|---:|---:|---:|---:|", + ] + ) + for row in sorted(rows, key=lambda item: str(item[0] or "")): + symbol, issuer, security_class, cusip, option_type = row[:5] + previous_shares, current_shares, change, change_pct = row[5:9] + lines.append( + f"| {_cell(symbol)} | {_cell(issuer)} | {_cell(security_class)} " + f"| {_cell(cusip)} | {_cell(option_type)} | {_fmt_shares(previous_shares)} " + f"| {_fmt_shares(current_shares)} | {_fmt_shares(change)} " + f"| {_fmt_percent(change_pct)} |" + ) + + lines.extend( + [ + "", + "## Filing History", + "", + "| Period | Holdings | Filing Type | Filed | SEC Accession |", + "|---|---:|---|---|---|", + ] + ) + for filing in filings[:20]: + lines.append( + f"| {filing.period} | {filing.holdings} | {filing.form_type} | {_iso_date(filing.filed)} " + f"| {_format_accession(filing.filing_id)} |" + ) return "\n".join(lines) -# --------------------------------------------------------------------------- -# Output: Cross-reference (manager × stock history) -# --------------------------------------------------------------------------- - - def _build_manager_stock_history( cik: str, cusip: str, manager_name: str, ticker: str ) -> str | None: - """Build Markdown for a specific manager's position history for a stock. - - Returns the Markdown string, or None if no data is available. - """ - data = _manager_cusip_history(cik, cusip) - if not data or not data.get("data"): - c.log(f"No position history found for {manager_name} in {ticker}.") + entries = _manager_cusip_history(cik, cusip).get("data", []) + if not entries: return None - - entries = data["data"] - - lines = [] - lines.append(f"# Position History: {manager_name} → {ticker}") - lines.append("") - lines.append(f"- **Manager CIK:** {cik}") - lines.append(f"- **CUSIP:** {cusip}") - lines.append(f"- **Source:** {_BASE}/manager/{cik}/cusip/{cusip}") - lines.append("") - - # Each entry: [[date, filing_slug], value, pct, shares, principal, date_filed, [year, quarter]] - lines.append("| Period | Shares | % of Portfolio | Filed |") - lines.append("|--------|--------|---------------|-------|") - - for e in entries: - pct = e[2] - shares = e[3] - date_filed = e[5] - period_info = e[6] # [year, quarter] - - period = f"Q{period_info[1]} {period_info[0]}" - lines.append(f"| {period} | {_fmt_shares(shares)} | {_fmt_pct(pct)} | {date_filed} |") - - return "\n".join(lines) - - -# --------------------------------------------------------------------------- -# Determine latest quarter -# --------------------------------------------------------------------------- - - -def _latest_quarter() -> tuple[int, int]: - """Return the most recent likely 13F quarter with data available. - - 13F filings are due 45 days after quarter end, and most large filers - submit within 45-60 days. We use a 50-day buffer from the *end* of - the quarter to be safe. Calendar quarters end Mar 31, Jun 30, Sep 30, - Dec 31. - """ - from datetime import date, timedelta - - today = date.today() - # Quarter end dates for the current and previous year - quarters = [] - for y in [today.year, today.year - 1]: - quarters.extend( - [ - (date(y, 3, 31), y, 1), - (date(y, 6, 30), y, 2), - (date(y, 9, 30), y, 3), - (date(y, 12, 31), y, 4), - ] + lines = [ + f"# Position History: {manager_name} → {ticker}", + "", + f"- **Manager CIK:** {cik}", + f"- **CUSIP:** {cusip}", + "- **Underlying records:** SEC Form 13F filings", + "", + "| Period | Shares | Filed | SEC Accession |", + "|---|---:|---|---|", + ] + for entry in entries: + period_info = entry[6] + filing_info = entry[0] + filing_slug = ( + filing_info[1] if isinstance(filing_info, list) and len(filing_info) > 1 else "" ) - # Sort descending and find the most recent quarter that ended 50+ days ago - quarters.sort(key=lambda x: x[0], reverse=True) - for end_date, year, q in quarters: - if today - end_date >= timedelta(days=50): - return (year, q) - # Fallback - return (today.year - 1, 4) + lines.append( + f"| Q{period_info[1]} {period_info[0]} | {_fmt_shares(entry[3])} | {entry[5]} " + f"| {_format_accession(filing_slug)} |" + ) + return "\n".join(lines) -# --------------------------------------------------------------------------- -# Main -# --------------------------------------------------------------------------- +def _write_or_fail(path, markdown: str | None, description: str) -> None: + if not markdown: + c.log(f"ERROR: no {description} data could be extracted.") + sys.exit(1) + c.emit(c.write_text(path, markdown)) -def main(): - p = argparse.ArgumentParser(description="Fetch 13F institutional holder data from 13f.info.") - # Stock-centric - p.add_argument("--ticker", help="Stock ticker to look up holders for.") - p.add_argument( - "--history", - action="store_true", - help="Show quarterly holder/share-count history (with --ticker).", - ) - p.add_argument( - "--year", type=int, default=None, help="Specific year for holder lookup (default: latest)." +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--ticker", help="Stock ticker for holder or manager-position queries.") + parser.add_argument("--manager", help="Manager name or exact 10-digit CIK.") + parser.add_argument( + "--history", action="store_true", help="Show history for a stock or manager." ) - p.add_argument( - "--quarter", - type=int, - default=None, - choices=[1, 2, 3, 4], - help="Specific quarter (1-4) for holder lookup.", - ) - p.add_argument( - "--top", type=int, default=25, help="Number of top holders to show (default: 25)." - ) - - # Manager-centric - p.add_argument("--manager", help="Manager name to search for.") - p.add_argument("--cik", help="Manager CIK (10-digit, e.g. 0001067983).") - - # Cross-reference - p.add_argument("--cusip", help="CUSIP for cross-reference with --cik.") + parser.add_argument("--year", type=int, help="Specific 13F reporting year.") + parser.add_argument("--quarter", type=int, choices=[1, 2, 3, 4], help="Specific quarter.") + parser.add_argument("--top", type=int, help="Top stock holders to show (default: 25).") + c.add_cache_arg(parser) + args = parser.parse_args() + + if not args.ticker and not args.manager: + parser.error("provide --ticker, --manager, or both") + if (args.year is None) != (args.quarter is None): + parser.error("--year and --quarter must be supplied together") + if args.top is not None and args.top <= 0: + parser.error("--top must be positive") + if args.manager and args.top is not None: + parser.error("--top applies only to stock-holder queries") + if args.manager and args.ticker and args.history: + parser.error("a manager-plus-ticker query already returns position history") - c.add_cache_arg(p) - args = p.parse_args() cache = c.cache_root(args.cache_dir) - - # Validate: at least one of --ticker, --manager, or --cik is required - if not args.ticker and not args.manager and not args.cik: - p.error("At least one of --ticker, --manager, or --cik is required.") - - # Mode 1: Cross-reference (--cik + --cusip) - if args.cik and args.cusip: - manager_name = args.manager or args.cik - ticker_label = args.ticker or args.cusip - md = _build_manager_stock_history(args.cik, args.cusip, manager_name, ticker_label) - if md: - out_dir = cache / "managers" - out_dir.mkdir(parents=True, exist_ok=True) - slug = c.safe_component(args.cik) - cusip_slug = c.safe_component(args.cusip) - path = c.write_text(out_dir / f"13f-xref_{slug}_{cusip_slug}.md", md) - c.emit(path) - return - - # Mode 2: Stock-centric (--ticker) - if args.ticker: - c.log(f"Resolving CUSIP for {args.ticker}...") - resolved = _resolve_cusip(args.ticker) - if not resolved: - c.log(f"ERROR: could not resolve CUSIP for {args.ticker}.") + try: + manager_data = _resolve_manager(args.manager) if args.manager else None + if args.manager and not manager_data: + c.log(f"ERROR: could not resolve manager '{args.manager}'.") sys.exit(1) - cusip, sym, issuer = resolved - c.log(f" {sym} → {issuer} (CUSIP: {cusip})") - if args.history: - md = _build_stock_history(args.ticker, cusip, issuer) - if md: - out_dir = c.company_dir(cache, None, ticker_hint=args.ticker) - path = c.write_text(out_dir / f"13f-history_{args.ticker.upper()}.md", md) - c.emit(path) + if args.manager and args.ticker: + cik, manager_name, _, _ = manager_data + resolved = _resolve_cusip(args.ticker) + if not resolved: + c.log(f"ERROR: could not resolve an exact CUSIP for ticker '{args.ticker}'.") + sys.exit(1) + cusip, _, _ = resolved + markdown = _build_manager_stock_history(cik, cusip, manager_name, args.ticker.upper()) + out_dir = c.company_dir(cache, None, ticker_hint=args.ticker) + _write_or_fail( + out_dir / f"13f-position-history_{c.safe_component(cik)}.md", + markdown, + "manager-position history", + ) return - # If --cik is also provided, show cross-reference - if args.cik: - manager_name = args.manager or args.cik - md = _build_manager_stock_history(args.cik, cusip, manager_name, args.ticker) - if md: - out_dir = c.company_dir(cache, None, ticker_hint=args.ticker) - slug = c.safe_component(args.cik) - path = c.write_text(out_dir / f"13f-xref_{slug}.md", md) - c.emit(path) + if args.ticker: + resolved = _resolve_cusip(args.ticker) + if not resolved: + c.log(f"ERROR: could not resolve an exact CUSIP for ticker '{args.ticker}'.") + sys.exit(1) + cusip, _symbol, issuer = resolved + out_dir = c.company_dir(cache, None, ticker_hint=args.ticker) + if args.history: + _write_or_fail( + out_dir / f"13f-history_{args.ticker.upper()}.md", + _build_stock_history(args.ticker.upper(), cusip, issuer), + "stock-holder history", + ) + return + year, quarter = ( + (args.year, args.quarter) if args.year is not None else _latest_quarter() + ) + _write_or_fail( + out_dir / f"13f-holders_{year}-Q{quarter}.md", + _build_stock_holders( + args.ticker.upper(), cusip, issuer, year, quarter, args.top or 25 + ), + f"Q{quarter} {year} holder", + ) return - year = args.year - quarter = args.quarter - if not year or not quarter: - y, q = _latest_quarter() - year = year or y - quarter = quarter or q - - c.log(f"Fetching Q{quarter} {year} holders for {sym}...") - md = _build_stock_holders(args.ticker, cusip, issuer, year, quarter, args.top) - if md: - out_dir = c.company_dir(cache, None, ticker_hint=args.ticker) - path = c.write_text(out_dir / f"13f-holders_{year}-Q{quarter}.md", md) - c.emit(path) - return - - # Mode 3: Manager-centric (--manager or --cik without --cusip) - if args.manager: - c.log(f"Searching for manager: {args.manager}") - resolved = _resolve_manager(args.manager) - if not resolved: - c.log(f"ERROR: could not find manager '{args.manager}'.") + cik, manager_name, _, manager_html = manager_data + filings = _parse_manager_filings(manager_html) + out_dir = cache / "managers" + out_dir.mkdir(parents=True, exist_ok=True) + if args.history: + _write_or_fail( + out_dir / f"13f-manager-history_{c.safe_component(cik)}.md", + _build_manager_history(cik, manager_name, filings), + "manager filing-history", + ) + return + selected = _select_manager_filing(filings, args.year, args.quarter) + if not selected: + period = f"Q{args.quarter} {args.year}" if args.year else "latest period" + c.log(f"ERROR: no complete manager portfolio found for {period}.") sys.exit(1) - cik, name = resolved - c.log(f" Found: {name} (CIK: {cik})") - md = _build_manager_holdings(cik, name) - if md: - out_dir = cache / "managers" - out_dir.mkdir(parents=True, exist_ok=True) - slug = c.safe_component(cik) - path = c.write_text(out_dir / f"13f-manager_{slug}.md", md) - c.emit(path) - return - - if args.cik: - md = _build_manager_holdings(args.cik, args.cik) - if md: - out_dir = cache / "managers" - out_dir.mkdir(parents=True, exist_ok=True) - slug = c.safe_component(args.cik) - path = c.write_text(out_dir / f"13f-manager_{slug}.md", md) - c.emit(path) - return + _write_or_fail( + out_dir / f"13f-manager_{c.safe_component(cik)}_{selected.period.replace(' ', '-')}.md", + _build_manager_holdings(cik, manager_name, selected), + "manager holding", + ) + except FetchFailure as exc: + c.log(f"ERROR: {exc}") + sys.exit(1) if __name__ == "__main__": diff --git a/skills/sec-edgar-skill/scripts/fetch_filing.py b/skills/sec-edgar-skill/scripts/fetch_filing.py index dffe814..2be30dd 100644 --- a/skills/sec-edgar-skill/scripts/fetch_filing.py +++ b/skills/sec-edgar-skill/scripts/fetch_filing.py @@ -1,218 +1,260 @@ -"""Fetch a single SEC filing (or one section / attachment) as Markdown. - -Writes into the cache (``//__[__].md``) -and prints the absolute path(s) to stdout. The script owns the filename so the -result is deterministic and re-discoverable; run with ``--help`` for all flags. -""" - -import argparse -import sys - -import _common as c - - -def _attachment_filename(stem: str, document: str, fallback: str) -> str: - name = c.safe_component(document or fallback) - return f"{stem}__{name if name.endswith('.md') else name + '.md'}" - - -def main(): - p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) - p.add_argument("--ticker", required=True, help="Ticker, CIK, or company name.") - p.add_argument( - "--form", required=True, help="Form type, e.g. 10-K, 10-Q, 8-K, 20-F, 40-F, 6-K, DEF 14A." - ) - p.add_argument("--year", type=int, help="Calendar year of the filing.") - p.add_argument("--quarter", type=int, choices=[1, 2, 3, 4], help="Calendar quarter.") - p.add_argument( - "--date", - help="Target a specific filing date (YYYY-MM-DD). " - "Picks the filing whose filing_date matches exactly, or the " - "most recent filing on or before that date.", - ) - p.add_argument( - "--section", - help='Extract one item by its SEC code, e.g. "Item 1A" ' - '(10-K/8-K/20-F) or "Part II, Item 1A" (10-Q). Pass "list" ' - "to print the item codes this filing actually contains.", - ) - p.add_argument( - "--attachment", - help='Attachment selector: "list" (print available), "all", an ' - "integer index, or a substring of the document/description " - '(e.g. "ex-99.1").', - ) - c.add_identity_arg(p) - c.add_cache_arg(p) - args = p.parse_args() - - c.resolve_identity(args.identity) - company = c.resolve_company(args.ticker) - - # amendments=False: the edgartools API natively excludes amended forms - # (10-K/A, 10-Q/A, etc.) when this is set. Amendments typically contain - # only the amended items (e.g. Part III), not the full filing, so picking - # one silently instead of the original loses most of the content. This - # matches guide_financials.md's amendments=False recommendation. - kwargs = {"form": args.form, "amendments": False} - if args.year: - kwargs["year"] = args.year - if args.quarter: - kwargs["quarter"] = args.quarter - try: - filings = list(company.get_filings(**kwargs)) - except Exception as exc: - c.log(f"ERROR: failed to list {args.form} filings: {exc}") - sys.exit(1) - if not filings: - c.log( - f"ERROR: no {args.form} filings for {args.ticker} " - f"(year={args.year}, quarter={args.quarter})." - ) - sys.exit(1) - - filings.sort(key=lambda f: f.filing_date, reverse=True) - - # --date: pick the filing whose date matches exactly, or the most recent - # filing on or before the given date. Useful for targeting a specific - # 8-K filed today without knowing its accession number. - if args.date: - from datetime import date as _date - - try: - _date.fromisoformat(args.date) - except ValueError: - c.log(f"ERROR: --date must be YYYY-MM-DD, got '{args.date}'.") - sys.exit(1) - candidates = [f for f in filings if str(getattr(f, "filing_date", "")) <= args.date] - if not candidates: - c.log(f"ERROR: no {args.form} filings for {args.ticker} on or before {args.date}.") - sys.exit(1) - # candidates is already sorted newest-first; pick the closest one. - filing = candidates[0] - c.log(f"--date {args.date}: selected filing dated {filing.filing_date}.") - else: - filing = filings[0] - - c.log(f"Resolved {filing.form} filed {filing.filing_date} (accession {filing.accession_no}).") - - out_dir = c.company_dir(c.cache_root(args.cache_dir), company, ticker_hint=args.ticker) - stem = c.filing_stem(filing) - - if args.attachment: - attachments = list(filing.attachments) - if not attachments: - c.log("ERROR: this filing has no attachments/exhibits.") - sys.exit(1) - - selector = args.attachment.lower() - if selector == "list": - c.log(f"{len(attachments)} attachment(s): index\tdocument\tdescription") - for i, att in enumerate(attachments): - print( - f"{i}\t{getattr(att, 'document', '') or ''}" - f"\t{getattr(att, 'description', '') or ''}" - ) - return - - if selector == "all": - saved = [] - for i, att in enumerate(attachments): - try: - fname = _attachment_filename( - stem, getattr(att, "document", ""), f"attachment_{i}" - ) - saved.append(c.write_text(out_dir / fname, att.markdown())) - c.log(f" saved {fname}") - except Exception as exc: - c.log(f" WARNING: attachment {i} failed: {exc}") - if not saved: - c.log("ERROR: no attachments could be converted.") - sys.exit(1) - for path in saved: - c.emit(path) - return - - selected = None - if args.attachment.isdigit(): - idx = int(args.attachment) - if not 0 <= idx < len(attachments): - c.log(f"ERROR: attachment index {idx} out of range (0-{len(attachments) - 1}).") - sys.exit(1) - selected = attachments[idx] - else: - needle = selector - for att in attachments: - doc = (getattr(att, "document", "") or "").lower() - desc = (getattr(att, "description", "") or "").lower() - if needle in doc or needle in desc: - selected = att - break - if selected is None: - c.log(f"ERROR: no attachment matched '{args.attachment}'.") - sys.exit(1) - - content = selected.markdown() - if not content: - c.log("ERROR: selected attachment produced no text.") - sys.exit(1) - fname = _attachment_filename(stem, getattr(selected, "document", ""), "attachment") - c.emit(c.write_text(out_dir / fname, content)) - return - - if args.section: - # Sections are addressed through the parsed data object, not markdown(): - # Filing.markdown() takes no section argument, so passing one there is - # silently ignored and returns the whole filing. obj()[code] slices it. - try: - obj = filing.obj() - except Exception as exc: - c.log( - f"ERROR: could not parse {filing.form} into a structured object " - f"to address by item: {exc}" - ) - sys.exit(1) - items = list(getattr(obj, "items", None) or []) - addressable = bool(items) and hasattr(obj, "__getitem__") - - if args.section.strip().lower() == "list": - if not addressable: - c.log( - f"{filing.form} is not item-addressable (it exposes no SEC " - f"item codes). Fetch the full filing (omit --section) and map " - f"it with list_headings.py." - ) - return - c.log(f"{filing.form} contains {len(items)} item(s):") - for it in items: - print(it) - return - - if not addressable: - c.log( - f"ERROR: {filing.form} is not item-addressable (no SEC item codes). " - f"Fetch the full filing (omit --section) and map it with " - f"list_headings.py, or pull an exhibit with --attachment." - ) - sys.exit(1) - - content = obj[args.section] - if content is None or not str(content).strip(): - c.log( - f"ERROR: item '{args.section}' is not present in this filing. " - f"Available items: {', '.join(items)}" - ) - sys.exit(1) - suffix = c.safe_component(args.section).lower() - c.emit(c.write_text(out_dir / f"{stem}__{suffix}.md", str(content))) - return - - content = filing.markdown() - if not content: - c.log("ERROR: filing produced no markdown (it may be exhibit-only; try --attachment list).") - sys.exit(1) - c.emit(c.write_text(out_dir / f"{stem}.md", content)) - - -if __name__ == "__main__": - main() +"""Fetch one SEC filing, section, or attachment as Markdown. + +Select either by accession, or by company plus form and optional period. Artifacts +are written to ``//__[__].md``; +run with ``--help`` for the complete selector contract. +""" + +from __future__ import annotations + +import argparse +import sys +from datetime import date +from pathlib import Path + +import _common as c + + +def _attachment_filename(stem: str, document: str, fallback: str) -> str: + name = c.safe_component(document or fallback) + return f"{stem}__{name if name.endswith('.md') else name + '.md'}" + + +def _validate_args(parser: argparse.ArgumentParser, args: argparse.Namespace) -> None: + if args.accession: + conflicting = [ + flag + for flag, value in ( + ("--ticker", args.ticker), + ("--form", args.form), + ("--year", args.year), + ("--quarter", args.quarter), + ("--on-or-before", args.on_or_before), + ) + if value is not None + ] + if conflicting: + parser.error(f"--accession cannot be combined with {', '.join(conflicting)}") + elif not args.ticker or not args.form: + parser.error("provide --accession, or provide both --ticker and --form") + + if args.attachment and args.section: + parser.error("--attachment and --section select different document types; choose one") + + +def _select_by_company(args: argparse.Namespace): + company = c.resolve_company(args.ticker) + kwargs: dict[str, object] = {"form": args.form, "amendments": False} + if args.year is not None: + kwargs["year"] = args.year + if args.quarter is not None: + kwargs["quarter"] = args.quarter + + try: + filings = list(company.get_filings(**kwargs)) + except Exception as exc: + c.log(f"ERROR: failed to list {args.form} filings: {exc}") + sys.exit(1) + if not filings: + c.log( + f"ERROR: no {args.form} filings for {args.ticker} " + f"(year={args.year}, quarter={args.quarter})." + ) + sys.exit(1) + + filings.sort(key=lambda filing: filing.filing_date, reverse=True) + if args.on_or_before: + try: + date.fromisoformat(args.on_or_before) + except ValueError: + c.log(f"ERROR: --on-or-before must be YYYY-MM-DD, got '{args.on_or_before}'.") + sys.exit(1) + filings = [ + filing + for filing in filings + if str(getattr(filing, "filing_date", "")) <= args.on_or_before + ] + if not filings: + c.log( + f"ERROR: no {args.form} filings for {args.ticker} on or before {args.on_or_before}." + ) + sys.exit(1) + c.log( + f"--on-or-before {args.on_or_before}: selected filing dated {filings[0].filing_date}." + ) + return company, filings[0] + + +def _cache_hit(path: Path, *, force: bool) -> bool: + if not force and path.is_file() and path.stat().st_size > 0: + c.log(f"Using cached artifact: {path.resolve()}") + return True + return False + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument( + "--accession", + help="Exact SEC accession. Cannot be combined with company/period selectors.", + ) + parser.add_argument("--ticker", help="Ticker, CIK, or company name.") + parser.add_argument("--form", help="Form type, e.g. 10-K, 10-Q, 8-K, 20-F, 40-F, 6-K, DEF 14A.") + parser.add_argument("--year", type=int, help="Calendar year of the filing.") + parser.add_argument("--quarter", type=int, choices=[1, 2, 3, 4], help="Calendar quarter.") + parser.add_argument( + "--on-or-before", + help="Select the most recent matching filing on or before YYYY-MM-DD.", + ) + parser.add_argument( + "--section", + help='Extract an SEC item code, or pass "list" to print available item codes.', + ) + parser.add_argument( + "--attachment", + help='Attachment selector: "list", "all", zero-based index, or document/description text.', + ) + c.add_identity_arg(parser) + c.add_cache_arg(parser) + c.add_force_arg(parser) + args = parser.parse_args() + _validate_args(parser, args) + + c.resolve_identity(args.identity) + if args.accession: + filing = c.resolve_filing(args.accession) + company = c.company_for_filing(filing) + else: + company, filing = _select_by_company(args) + + accession = getattr(filing, "accession_no", "") or getattr(filing, "accession_number", "") + c.log(f"Resolved {filing.form} filed {filing.filing_date} (accession {accession}).") + + out_dir = c.company_dir( + c.cache_root(args.cache_dir), + company, + ticker_hint=args.ticker if not args.accession else None, + ) + stem = c.filing_stem(filing) + + if args.attachment: + attachments = list(filing.attachments) + if not attachments: + c.log("ERROR: this filing has no attachments/exhibits.") + sys.exit(1) + + selector = args.attachment.lower() + if selector == "list": + c.log(f"{len(attachments)} attachment(s): index\tdocument\tdescription") + for index, attachment in enumerate(attachments): + print( + f"{index}\t{getattr(attachment, 'document', '') or ''}" + f"\t{getattr(attachment, 'description', '') or ''}" + ) + return + + if selector == "all": + saved: list[Path] = [] + for index, attachment in enumerate(attachments): + filename = _attachment_filename( + stem, getattr(attachment, "document", ""), f"attachment_{index}" + ) + path = out_dir / filename + if _cache_hit(path, force=args.force): + saved.append(path) + continue + try: + content = attachment.markdown() + if not content: + raise ValueError("attachment produced no text") + saved.append(c.write_text(path, content)) + c.log(f" saved {filename}") + except Exception as exc: + c.log(f" WARNING: attachment {index} failed: {exc}") + if not saved: + c.log("ERROR: no attachments could be converted.") + sys.exit(1) + for path in saved: + c.emit(path) + return + + selected = None + if args.attachment.isdigit(): + index = int(args.attachment) + if not 0 <= index < len(attachments): + c.log(f"ERROR: attachment index {index} out of range (0-{len(attachments) - 1}).") + sys.exit(1) + selected = attachments[index] + else: + for attachment in attachments: + document = (getattr(attachment, "document", "") or "").lower() + description = (getattr(attachment, "description", "") or "").lower() + if selector in document or selector in description: + selected = attachment + break + if selected is None: + c.log(f"ERROR: no attachment matched '{args.attachment}'.") + sys.exit(1) + + filename = _attachment_filename(stem, getattr(selected, "document", ""), "attachment") + path = out_dir / filename + if c.emit_cached(path, force=args.force): + return + content = selected.markdown() + if not content: + c.log("ERROR: selected attachment produced no text.") + sys.exit(1) + c.emit(c.write_text(path, content)) + return + + if args.section: + if args.section.strip().lower() != "list": + suffix = c.safe_component(args.section).lower() + path = out_dir / f"{stem}__{suffix}.md" + if c.emit_cached(path, force=args.force): + return + try: + obj = filing.obj() + except Exception as exc: + c.log(f"ERROR: could not parse {filing.form} into an item-addressable object: {exc}") + sys.exit(1) + items = list(getattr(obj, "items", None) or []) + addressable = bool(items) and hasattr(obj, "__getitem__") + + if args.section.strip().lower() == "list": + if not addressable: + c.log( + f"{filing.form} is not item-addressable. Fetch the full filing and " + "navigate its contents or attachments." + ) + return + c.log(f"{filing.form} contains {len(items)} item(s):") + for item in items: + print(item) + return + + if not addressable: + c.log( + f"ERROR: {filing.form} is not item-addressable. Fetch the full filing or " + "an attachment instead." + ) + sys.exit(1) + content = obj[args.section] + if content is None or not str(content).strip(): + c.log(f"ERROR: item '{args.section}' is absent. Available items: {', '.join(items)}") + sys.exit(1) + c.emit(c.write_text(path, str(content))) + return + + path = out_dir / f"{stem}.md" + if c.emit_cached(path, force=args.force): + return + content = filing.markdown() + if not content: + c.log("ERROR: filing produced no Markdown; it may be exhibit-only. Try --attachment list.") + sys.exit(1) + c.emit(c.write_text(path, content)) + + +if __name__ == "__main__": + main() diff --git a/skills/sec-edgar-skill/scripts/fetch_filings.py b/skills/sec-edgar-skill/scripts/fetch_filings.py index 59739fb..c25fa65 100644 --- a/skills/sec-edgar-skill/scripts/fetch_filings.py +++ b/skills/sec-edgar-skill/scripts/fetch_filings.py @@ -1,86 +1,123 @@ -"""Bulk-fetch SEC filings across a year range as Markdown into the cache. - -Each filing is written to ``//__.md`` and its -absolute path printed to stdout. With ``--attachments`` each filing's text -attachments are saved alongside it (e.g. a 6-K's Exhibit 99.1). Run ``--help`` -for all flags. -""" - -import argparse -import sys - -import _common as c - -_BINARY = (".jpg", ".jpeg", ".png", ".gif", ".zip", ".pdf", ".xlsx", ".xls") - - -def main(): - p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) - p.add_argument("--ticker", required=True, help="Ticker, CIK, or company name.") - p.add_argument("--form", required=True, help="Form type, e.g. 10-Q, 8-K, 6-K.") - p.add_argument("--start-year", type=int, help="First calendar year (inclusive).") - p.add_argument("--end-year", type=int, help="Last calendar year (inclusive).") - p.add_argument( - "--attachments", action="store_true", help="Also save each filing's text attachments." - ) - c.add_identity_arg(p) - c.add_cache_arg(p) - args = p.parse_args() - - c.resolve_identity(args.identity) - company = c.resolve_company(args.ticker) - - kwargs = {"form": args.form} - if args.start_year or args.end_year: - start = f"{args.start_year}-01-01" if args.start_year else "1994-01-01" - end = f"{args.end_year}-12-31" if args.end_year else "2100-12-31" - kwargs["date"] = f"{start}:{end}" - try: - filings = list(company.get_filings(**kwargs)) - except Exception as exc: - c.log(f"ERROR: failed to list {args.form} filings: {exc}") - sys.exit(1) - if not filings: - c.log("No filings matched the criteria.") - return - - out_dir = c.company_dir(c.cache_root(args.cache_dir), company, ticker_hint=args.ticker) - c.log(f"Found {len(filings)} {args.form} filing(s) -> {out_dir}") - - saved = [] - for i, filing in enumerate(filings, 1): - stem = c.filing_stem(filing) - c.log(f"[{i}/{len(filings)}] {filing.form} {filing.filing_date}") - try: - content = filing.markdown() - if content: - saved.append(c.write_text(out_dir / f"{stem}.md", content)) - else: - c.log(" (empty body — likely exhibit-only; use --attachments)") - except Exception as exc: - c.log(f" ERROR main body: {exc}") - - if args.attachments: - try: - for j, att in enumerate(list(filing.attachments)): - doc = getattr(att, "document", "") or f"attachment_{j}" - if doc.lower().endswith(_BINARY): - continue - try: - name = c.safe_component(doc) - name = name if name.endswith(".md") else name + ".md" - c.write_text(out_dir / f"{stem}__{name}", att.markdown()) - except Exception: - continue - except Exception as exc: - c.log(f" WARNING attachments: {exc}") - - if not saved: - c.log("ERROR: nothing could be saved.") - sys.exit(1) - for path in saved: - c.emit(path) - - -if __name__ == "__main__": - main() +"""Bulk-fetch SEC filings and attachments across a filing-date range. + +Bulk retrieval intentionally includes both original filings and amendments. Each +saved or reused artifact is emitted as an absolute path; run with ``--help`` for +all selectors. +""" + +from __future__ import annotations + +import argparse +import sys +from pathlib import Path + +import _common as c + +_BINARY = (".jpg", ".jpeg", ".png", ".gif", ".zip", ".pdf", ".xlsx", ".xls") + + +def _cached(path: Path, *, force: bool) -> bool: + return not force and path.is_file() and path.stat().st_size > 0 + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--ticker", required=True, help="Ticker, CIK, or company name.") + parser.add_argument("--form", required=True, help="Form type, e.g. 10-Q, 8-K, 6-K.") + parser.add_argument("--start-year", type=int, help="First calendar year (inclusive).") + parser.add_argument("--end-year", type=int, help="Last calendar year (inclusive).") + parser.add_argument( + "--attachments", action="store_true", help="Also save each filing's text attachments." + ) + c.add_identity_arg(parser) + c.add_cache_arg(parser) + c.add_force_arg(parser) + args = parser.parse_args() + + if args.start_year and args.end_year and args.start_year > args.end_year: + parser.error("--start-year cannot be later than --end-year") + + c.resolve_identity(args.identity) + company = c.resolve_company(args.ticker) + + # A bulk archive should preserve the full record: original and amended forms. + kwargs: dict[str, object] = {"form": args.form, "amendments": True} + if args.start_year or args.end_year: + start = f"{args.start_year}-01-01" if args.start_year else "1994-01-01" + end = f"{args.end_year}-12-31" if args.end_year else "2100-12-31" + kwargs["date"] = f"{start}:{end}" + try: + filings = list(company.get_filings(**kwargs)) + except Exception as exc: + c.log(f"ERROR: failed to list {args.form} filings: {exc}") + sys.exit(1) + if not filings: + c.log("ERROR: no filings matched the requested form and date range.") + sys.exit(1) + + out_dir = c.company_dir(c.cache_root(args.cache_dir), company, ticker_hint=args.ticker) + c.log( + f"Found {len(filings)} {args.form} filing(s), including amendments when present, " + f"-> {out_dir}" + ) + + saved: list[Path] = [] + failures = 0 + for index, filing in enumerate(filings, 1): + stem = c.filing_stem(filing) + c.log(f"[{index}/{len(filings)}] {filing.form} {filing.filing_date}") + body_path = out_dir / f"{stem}.md" + if _cached(body_path, force=args.force): + c.log(f" using cached {body_path.name}") + saved.append(body_path) + else: + try: + content = filing.markdown() + if content: + saved.append(c.write_text(body_path, content)) + else: + c.log(" empty body — likely exhibit-only; use --attachments") + except Exception as exc: + failures += 1 + c.log(f" WARNING: main body failed: {exc}") + + if not args.attachments: + continue + try: + attachments = list(filing.attachments) + except Exception as exc: + failures += 1 + c.log(f" WARNING: could not list attachments: {exc}") + continue + + for attachment_index, attachment in enumerate(attachments): + document = getattr(attachment, "document", "") or f"attachment_{attachment_index}" + if document.lower().endswith(_BINARY): + continue + name = c.safe_component(document) + name = name if name.endswith(".md") else name + ".md" + attachment_path = out_dir / f"{stem}__{name}" + if _cached(attachment_path, force=args.force): + c.log(f" using cached {attachment_path.name}") + saved.append(attachment_path) + continue + try: + content = attachment.markdown() + if not content: + raise ValueError("attachment produced no text") + saved.append(c.write_text(attachment_path, content)) + except Exception as exc: + failures += 1 + c.log(f" WARNING: attachment {attachment_index} ({document}) failed: {exc}") + + if not saved: + c.log("ERROR: no filing bodies or attachments could be saved.") + sys.exit(1) + if failures: + c.log(f"WARNING: completed with {failures} failed conversion(s).") + for path in dict.fromkeys(saved): + c.emit(path) + + +if __name__ == "__main__": + main() diff --git a/skills/sec-edgar-skill/scripts/fetch_insider_trades.py b/skills/sec-edgar-skill/scripts/fetch_insider_trades.py index 0603cb5..19f822a 100644 --- a/skills/sec-edgar-skill/scripts/fetch_insider_trades.py +++ b/skills/sec-edgar-skill/scripts/fetch_insider_trades.py @@ -1,40 +1,23 @@ -"""Fetch insider transactions (Form 3/4/5) for a single company. +"""Fetch one company's Form 4 insider transactions. -Pulls Form 4 filings for a ticker within a date range and extracts open-market -purchases (code P), sales (code S), option exercises (code M), and tax -withholdings (code F). Writes a Markdown summary into the cache and emits -the path to stdout. - -This is a **per-company** tool for due-diligence — "what are insiders of -company X doing?" — not a market-wide scanner. For broad market scanning -of insider clusters and rip/dip buys, see signal-sweep's scan_insiders.py. - -Usage: - # Last 12 months (default) - python scripts/fetch_insider_trades.py --ticker AAPL - - # Custom date range - python scripts/fetch_insider_trades.py --ticker AAPL --start 2025-01-01 --end 2026-06-17 - - # Only open-market purchases - python scripts/fetch_insider_trades.py --ticker AAPL --start 2025-06-01 --buys-only - -Output is written to ``//insider-trades__.md`` -(or ``insider-buys_...`` with --buys-only). The absolute path is emitted to -stdout; progress goes to stderr. +The report distinguishes transaction dates from filing dates and includes only +successfully parsed filings. Valid no-data results are emitted; retrieval failure +or a period in which every filing failed to parse exits nonzero. """ from __future__ import annotations import argparse import sys -from datetime import date, timedelta +from dataclasses import dataclass +from datetime import date, datetime, timedelta from pathlib import Path +import pandas as pd + sys.path.insert(0, str(Path(__file__).resolve().parent)) import _common as c -# Transaction code labels _CODE_LABELS = { "P": "Purchase", "S": "Sale", @@ -48,69 +31,113 @@ } +class FetchFailure(RuntimeError): + """Raised when a Form 4 query cannot be completed.""" + + +@dataclass +class FetchResult: + """Transactions and completeness metadata for one Form 4 query.""" + + transactions: list[dict] + filings_found: int + filings_parsed: int + filings_failed: int + + +def _missing(value: object) -> bool: + try: + result = pd.isna(value) + return bool(result) if not hasattr(result, "all") else bool(result.all()) + except (TypeError, ValueError): + return value is None + + +def _number(value: object) -> float | None: + if _missing(value): + return None + try: + return float(value) + except (TypeError, ValueError): + return None + + +def _cell(value: object) -> str: + if _missing(value): + return "" + return str(value).replace("|", "\\|").replace("\n", " ").strip() + + +def _date_text(value: object) -> str: + if _missing(value): + return "n/a" + if hasattr(value, "date"): + try: + return value.date().isoformat() + except (AttributeError, TypeError, ValueError): + pass + text = str(value) + return text[:10] if len(text) >= 10 else text + + def _code_label(code: str) -> str: return _CODE_LABELS.get(code, code or "Unknown") -def _fmt_shares(s) -> str: - if s is None: +def _fmt_shares(value: object) -> str: + number = _number(value) + if number is None: return "n/a" - s = float(s) - if abs(s) >= 1_000_000: - return f"{s / 1_000_000:,.2f}M" - if abs(s) >= 1_000: - return f"{s / 1_000:,.1f}K" - return f"{s:,.0f}" + if abs(number) >= 1_000_000: + return f"{number / 1_000_000:,.2f}M" + if abs(number) >= 1_000: + return f"{number / 1_000:,.1f}K" + return f"{number:,.0f}" -def _fmt_price(p) -> str: - if p is None or p == 0: +def _fmt_price(value: object) -> str: + number = _number(value) + if number in (None, 0): return "n/a" - return f"${float(p):,.2f}" + return f"${number:,.2f}" -def _fmt_value(shares, price) -> str: - if shares is None or price is None or price == 0: +def _fmt_value(shares: object, price: object) -> str: + share_number = _number(shares) + price_number = _number(price) + if share_number is None or price_number in (None, 0): return "n/a" - val = abs(float(shares) * float(price)) - if val >= 1_000_000: - return f"${val / 1_000_000:,.2f}M" - if val >= 1_000: - return f"${val / 1_000:,.1f}K" - return f"${val:,.0f}" - + value = abs(share_number * price_number) + if value >= 1_000_000: + return f"${value / 1_000_000:,.2f}M" + if value >= 1_000: + return f"${value / 1_000:,.1f}K" + return f"${value:,.0f}" -def _parse_date(s: str) -> date: - """Parse YYYY-MM-DD into a date object.""" - from datetime import datetime as dt - return dt.strptime(s, "%Y-%m-%d").date() +def _parse_date(value: str) -> date: + return datetime.strptime(value, "%Y-%m-%d").date() def fetch_insider_trades( ticker: str, start: date, end: date, buys_only: bool = False -) -> list[dict]: - """Fetch and parse Form 4 filings for a ticker within [start, end]. - - Returns a list of transaction dicts, newest first. - """ +) -> FetchResult: + """Fetch and parse original Form 4 filings whose filing dates fall in the range.""" company = c.resolve_company(ticker) date_range = f"{start.isoformat()}:{end.isoformat()}" c.log(f"Fetching Form 4 filings for {ticker} ({date_range})...") try: - form4s = company.get_filings(form="4", date=date_range) + form4s = company.get_filings(form="4", date=date_range, amendments=False) except Exception as exc: - c.log(f"ERROR: could not fetch Form 4 filings: {exc}") - return [] - - if form4s is None or len(form4s) == 0: - c.log(f"No Form 4 filings found for {ticker} in {date_range}.") - return [] + raise FetchFailure(f"could not fetch Form 4 filings: {exc}") from exc - total = len(form4s) - c.log(f" Found {total} Form 4 filings in range.") + total = len(form4s) if form4s is not None else 0 + if total == 0: + c.log(f"No original Form 4 filings found for {ticker} in {date_range}.") + return FetchResult([], 0, 0, 0) + c.log(f" Found {total} original Form 4 filings in range.") transactions = [] parsed = 0 errors = 0 @@ -118,185 +145,193 @@ def fetch_insider_trades( for filing in form4s: try: obj = filing.obj() - except Exception: - errors += 1 - continue - - filing_date = str(getattr(filing, "filing_date", "")) - insider_name = getattr(obj, "insider_name", None) or "Unknown" - position = getattr(obj, "position", None) or "Unknown" - - # Get transactions via to_dataframe() - try: - txn_df = obj.to_dataframe() - except Exception: + transaction_frame = obj.to_dataframe() + except Exception as exc: errors += 1 + c.log(f" WARNING: could not parse {getattr(filing, 'accession_no', '')}: {exc}") continue - if txn_df is None or len(txn_df) == 0: - parsed += 1 + parsed += 1 + if transaction_frame is None or len(transaction_frame) == 0: continue - for _, row in txn_df.iterrows(): - code = row.get("Code", "") - shares = row.get("Shares", 0) - price = row.get("Price", 0) - remaining = row.get("Remaining Shares", None) + filed_date = _date_text(getattr(filing, "filing_date", "")) + accession = str( + getattr(filing, "accession_no", "") or getattr(filing, "accession_number", "") + ) + default_name = getattr(obj, "insider_name", None) or "Unknown" + default_position = getattr(obj, "position", None) or "Unknown" + for _, row in transaction_frame.iterrows(): + code = _cell(row.get("Code", "")) if buys_only and code != "P": continue - transactions.append( { - "date": filing_date, - "insider": str(insider_name), - "role": str(position), - "code": str(code or ""), + "transaction_date": _date_text(row.get("Date")), + "filed_date": filed_date, + "insider": _cell(row.get("Insider", default_name)) or _cell(default_name), + "role": _cell(row.get("Position", default_position)) or _cell(default_position), + "code": code, "type": _code_label(code), - "shares": shares, - "price": price, - "remaining": remaining, - "accession": ( - getattr(filing, "accession_no", "") - or getattr(filing, "accession_number", "") - ), + "shares": _number(row.get("Shares")), + "price": _number(row.get("Price")), + "remaining": _number(row.get("Remaining Shares")), + "accession": accession, } ) - parsed += 1 if parsed % 25 == 0: c.log(f" Parsed {parsed}/{total} filings...") - if errors > 0: - c.log(f" ({errors} filings could not be parsed)") - - c.log(f" Parsed {parsed} Form 4 filings, found {len(transactions)} transactions.") - return transactions + transactions.sort( + key=lambda item: (item["transaction_date"], item["filed_date"], item["accession"]), + reverse=True, + ) + c.log(f" Parsed {parsed}/{total} filings and found {len(transactions)} transactions.") + return FetchResult(transactions, total, parsed, errors) def _build_markdown( - ticker: str, transactions: list[dict], start: date, end: date, buys_only: bool + ticker: str, result: FetchResult, start: date, end: date, buys_only: bool ) -> str: - """Build a Markdown summary of insider transactions.""" - lines = [] - lines.append(f"# Insider Transactions: {ticker}") - lines.append("") - lines.append(f"- **Period:** {start.isoformat()} to {end.isoformat()}") - lines.append("") + transactions = result.transactions + lines = [ + f"# Insider Transactions: {ticker}", + "", + f"- **Form 4 filing period queried:** {start.isoformat()} to {end.isoformat()}", + f"- **Filings found:** {result.filings_found}", + f"- **Filings parsed:** {result.filings_parsed}", + f"- **Filings failed:** {result.filings_failed}", + f"- **Completeness:** {'Partial' if result.filings_failed else 'Complete'}", + "", + ] if not transactions: - lines.append("No insider transactions found in this period.") + lines.append("No matching insider transactions were found in successfully parsed filings.") return "\n".join(lines) - # --- Aggregate stats --- - buys = [t for t in transactions if t["code"] == "P"] - sells = [t for t in transactions if t["code"] == "S"] - exercises = [t for t in transactions if t["code"] == "M"] + buys = [transaction for transaction in transactions if transaction["code"] == "P"] + sells = [transaction for transaction in transactions if transaction["code"] == "S"] + exercises = [transaction for transaction in transactions if transaction["code"] == "M"] buy_value = sum( - abs(t["shares"] * (t["price"] or 0)) for t in buys if t["shares"] and t["price"] + abs(transaction["shares"] * transaction["price"]) + for transaction in buys + if transaction["shares"] is not None and transaction["price"] is not None ) sell_value = sum( - abs(t["shares"] * (t["price"] or 0)) for t in sells if t["shares"] and t["price"] + abs(transaction["shares"] * transaction["price"]) + for transaction in sells + if transaction["shares"] is not None and transaction["price"] is not None ) - unique_buyers = set(t["insider"] for t in buys) - unique_sellers = set(t["insider"] for t in sells) - - lines.append("## Summary") - lines.append("") - lines.append("| Metric | Value |") - lines.append("|--------|-------|") - lines.append(f"| Open-market purchases | {len(buys)} |") - lines.append(f"| Unique buyers | {len(unique_buyers)} |") - lines.append(f"| Total buy value | {_fmt_value(1, buy_value) if buy_value else 'n/a'} |") + lines.extend( + [ + "## Summary", + "", + "| Metric | Value |", + "|--------|-------|", + f"| Open-market purchases | {len(buys)} |", + f"| Unique buyers | {len({transaction['insider'] for transaction in buys})} |", + f"| Total buy value | {_fmt_value(1, buy_value) if buy_value else 'n/a'} |", + ] + ) if not buys_only: - lines.append(f"| Open-market sales | {len(sells)} |") - lines.append(f"| Unique sellers | {len(unique_sellers)} |") - lines.append(f"| Total sell value | {_fmt_value(1, sell_value) if sell_value else 'n/a'} |") - lines.append(f"| Option exercises | {len(exercises)} |") - ratio = "n/a" - if sell_value > 0: - ratio = f"{buy_value / sell_value:.2f}x" - elif buy_value > 0: - ratio = "∞ (no sales)" - lines.append(f"| Buy/sell ratio (by $) | {ratio} |") - lines.append("") - - # --- Transaction detail table --- - label = "Open-Market Purchases" if buys_only else "All Transactions" - lines.append(f"## {label}") - lines.append("") - lines.append("| Date | Insider | Role | Type | Shares | Price | Value | Remaining |") - lines.append("|------|---------|------|------|--------|-------|-------|-----------|") - - for t in transactions: + lines.extend( + [ + f"| Open-market sales | {len(sells)} |", + f"| Unique sellers | {len({transaction['insider'] for transaction in sells})} |", + f"| Total sell value | {_fmt_value(1, sell_value) if sell_value else 'n/a'} |", + f"| Option exercises | {len(exercises)} |", + ] + ) + lines.extend( + [ + "", + "## " + ("Open-Market Purchases" if buys_only else "Transactions"), + "", + "| Transaction Date | Filed | Insider | Role | Type | Shares | Price | Value | Remaining | Accession |", + "|---|---|---|---|---|---:|---:|---:|---:|---|", + ] + ) + for transaction in transactions: lines.append( - f"| {t['date']} | {t['insider']} | {t['role']} " - f"| {t['type']} | {_fmt_shares(t['shares'])} " - f"| {_fmt_price(t['price'])} " - f"| {_fmt_value(t['shares'], t['price'])} " - f"| {_fmt_shares(t['remaining'])} |" + f"| {transaction['transaction_date']} | {transaction['filed_date']} " + f"| {transaction['insider']} | {transaction['role']} | {transaction['type']} " + f"| {_fmt_shares(transaction['shares'])} | {_fmt_price(transaction['price'])} " + f"| {_fmt_value(transaction['shares'], transaction['price'])} " + f"| {_fmt_shares(transaction['remaining'])} | {transaction['accession']} |" ) - lines.append("") - - # --- Insider-level summary (who is buying / selling) --- if buys: - lines.append("## Buyers") - lines.append("") - lines.append("| Insider | Role | Purchases | Total Shares | Total Value |") - lines.append("|---------|------|-----------|--------------|-------------|") + lines.extend( + [ + "", + "## Buyers", + "", + "| Insider | Role | Purchases | Total Shares | Total Value |", + "|---|---|---:|---:|---:|", + ] + ) buyer_data: dict[str, dict] = {} - for t in buys: - key = t["insider"] - if key not in buyer_data: - buyer_data[key] = {"role": t["role"], "count": 0, "shares": 0, "value": 0} - buyer_data[key]["count"] += 1 - buyer_data[key]["shares"] += abs(t["shares"] or 0) - buyer_data[key]["value"] += abs((t["shares"] or 0) * (t["price"] or 0)) - for name, d in sorted(buyer_data.items(), key=lambda x: x[1]["value"], reverse=True): + for transaction in buys: + key = transaction["insider"] + buyer = buyer_data.setdefault( + key, + {"role": transaction["role"], "count": 0, "shares": 0.0, "value": 0.0}, + ) + buyer["count"] += 1 + buyer["shares"] += abs(transaction["shares"] or 0) + buyer["value"] += abs((transaction["shares"] or 0) * (transaction["price"] or 0)) + for name, buyer in sorted( + buyer_data.items(), key=lambda item: item[1]["value"], reverse=True + ): lines.append( - f"| {name} | {d['role']} | {d['count']} " - f"| {_fmt_shares(d['shares'])} " - f"| {_fmt_value(1, d['value'])} |" + f"| {name} | {buyer['role']} | {buyer['count']} " + f"| {_fmt_shares(buyer['shares'])} | {_fmt_value(1, buyer['value'])} |" ) - lines.append("") + lines.append("") return "\n".join(lines) -def main(): - p = argparse.ArgumentParser(description="Fetch insider transactions (Form 4) for a company.") - p.add_argument("--ticker", required=True, help="Stock ticker to look up.") - p.add_argument("--start", help="Start date YYYY-MM-DD (default: 12 months ago).") - p.add_argument("--end", help="End date YYYY-MM-DD (default: today).") - p.add_argument( - "--buys-only", action="store_true", help="Show only open-market purchases (code P)." +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--ticker", required=True, help="Stock ticker to look up.") + parser.add_argument( + "--start", help="Form 4 filing-date start, YYYY-MM-DD (default: one year ago)." ) - c.add_identity_arg(p) - c.add_cache_arg(p) - args = p.parse_args() + parser.add_argument("--end", help="Form 4 filing-date end, YYYY-MM-DD (default: today).") + parser.add_argument( + "--buys-only", action="store_true", help="Keep only open-market purchases (code P)." + ) + c.add_identity_arg(parser) + c.add_cache_arg(parser) + args = parser.parse_args() c.resolve_identity(args.identity) - cache = c.cache_root(args.cache_dir) - - end = _parse_date(args.end) if args.end else date.today() - start = _parse_date(args.start) if args.start else end - timedelta(days=365) - + try: + end = _parse_date(args.end) if args.end else date.today() + start = _parse_date(args.start) if args.start else end - timedelta(days=365) + except ValueError as exc: + parser.error(f"dates must use YYYY-MM-DD: {exc}") if start > end: - c.log("ERROR: --start is after --end.") - sys.exit(1) - - transactions = fetch_insider_trades(args.ticker, start=start, end=end, buys_only=args.buys_only) + parser.error("--start cannot be later than --end") - md = _build_markdown(args.ticker, transactions, start, end, args.buys_only) + try: + result = fetch_insider_trades(args.ticker, start=start, end=end, buys_only=args.buys_only) + except FetchFailure as exc: + c.log(f"ERROR: {exc}") + sys.exit(1) + if result.filings_found and result.filings_parsed == 0: + c.log("ERROR: every Form 4 filing failed to parse; no report was written.") + sys.exit(1) - # Write to cache: //insider-trades__.md - out_dir = c.company_dir(cache, None, ticker_hint=args.ticker) + markdown = _build_markdown(args.ticker, result, start, end, args.buys_only) + out_dir = c.company_dir(c.cache_root(args.cache_dir), None, ticker_hint=args.ticker) prefix = "insider-buys" if args.buys_only else "insider-trades" - fname = f"{prefix}_{start.isoformat()}_{end.isoformat()}.md" - path = c.write_text(out_dir / fname, md) + path = c.write_text(out_dir / f"{prefix}_{start.isoformat()}_{end.isoformat()}.md", markdown) c.emit(path) diff --git a/skills/sec-edgar-skill/scripts/orient.py b/skills/sec-edgar-skill/scripts/orient.py index 4778c03..9c90ea9 100644 --- a/skills/sec-edgar-skill/scripts/orient.py +++ b/skills/sec-edgar-skill/scripts/orient.py @@ -1,90 +1,87 @@ -"""Orient on a company before extracting — the mandatory first step. - -Resolves the company, prints its ``.to_context()`` summary, surveys the mix of -forms it has actually filed over a recent window (per-form counts with date -ranges), and lists the most recent filings. Output is compact Markdown to stdout; -run with ``--help`` for all flags. - -Run this first for any filings work. It is the cheapest way to see what a company -actually files *now* and how that has changed over time, so you fetch the right -forms instead of assuming a form set from memory. The output is neutral — it shows -the filing history and leaves what is significant to you (or the framework driving -you) to decide. -""" - -import argparse -import datetime as _dt -import sys - -import _common as c - - -def main(): - p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) - p.add_argument("--ticker", required=True, help="Ticker, CIK, or company name.") - p.add_argument( - "--years", - type=int, - default=3, - help="Survey the filing mix over the last N calendar years (default: 3).", - ) - p.add_argument( - "--recent", type=int, default=15, help="Also list the N most recent filings (default: 15)." - ) - c.add_identity_arg(p) - args = p.parse_args() - - c.resolve_identity(args.identity) - company = c.resolve_company(args.ticker) - - # 1) Cheap metadata summary. - print(f"# Orientation: {args.ticker.upper()}\n") - try: - print(company.to_context()) - except Exception as exc: - c.log(f"WARNING: .to_context() unavailable ({exc}); printing basic metadata.") - print(f"COMPANY: {getattr(company, 'name', 'n/a')}") - print(f"CIK: {getattr(company, 'cik', 'n/a')}") - print(f"SIC: {getattr(company, 'sic', 'n/a')}") - - # 2) Filing-mix survey over the window. - this_year = _dt.date.today().year - start_year = this_year - max(args.years - 1, 0) - date_range = f"{start_year}-01-01:{this_year}-12-31" - try: - df = company.get_filings(date=date_range).to_pandas() - except Exception as exc: - c.log(f"ERROR: could not survey filings for {date_range}: {exc}") - sys.exit(1) - - print(f"\n## Filing mix {start_year}-{this_year}") - if df is None or len(df) == 0: - print("- (no filings in this window; widen --years)") - return - - df = df.copy() - df["_date"] = df["filing_date"].astype(str) - - rows = [] - for form, grp in df.groupby("form"): - d = grp["_date"] - rows.append((str(form), len(grp), d.min(), d.max())) - rows.sort(key=lambda r: (r[1], r[3]), reverse=True) - - print("\n| Form | Count | Earliest | Latest |") - print("| :-- | --: | :-- | :-- |") - for form, n, lo, hi in rows: - print(f"| {form} | {n} | {lo} | {hi} |") - - # 3) Most recent filings. - n_recent = min(args.recent, len(df)) - print(f"\n## {n_recent} most recent filings") - recent = df.sort_values("_date", ascending=False).head(n_recent) - print("\n| Date | Form | Accession |") - print("| :-- | :-- | :-- |") - for _, row in recent.iterrows(): - print(f"| {row['_date']} | {row['form']} | {row.get('accession_number', '')} |") - - -if __name__ == "__main__": - main() +"""Survey a company's identity, filing mix, and recent accessions. + +Use this when the relevant form or filing is not yet known. It resolves the +company, prints its compact ``.to_context()`` summary, tabulates forms over a +recent window, and lists recent filings. It does not download filing contents. +""" + +import argparse +import datetime as _dt +import sys + +import _common as c + + +def main(): + p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + p.add_argument("--ticker", required=True, help="Ticker, CIK, or company name.") + p.add_argument( + "--years", + type=int, + default=3, + help="Survey the filing mix over the last N calendar years (default: 3).", + ) + p.add_argument( + "--recent", type=int, default=15, help="Also list the N most recent filings (default: 15)." + ) + c.add_identity_arg(p) + args = p.parse_args() + if args.years <= 0: + p.error("--years must be positive") + if args.recent < 0: + p.error("--recent cannot be negative") + + c.resolve_identity(args.identity) + company = c.resolve_company(args.ticker) + + # 1) Cheap metadata summary. + print(f"# Orientation: {args.ticker.upper()}\n") + try: + print(company.to_context()) + except Exception as exc: + c.log(f"WARNING: .to_context() unavailable ({exc}); printing basic metadata.") + print(f"COMPANY: {getattr(company, 'name', 'n/a')}") + print(f"CIK: {getattr(company, 'cik', 'n/a')}") + print(f"SIC: {getattr(company, 'sic', 'n/a')}") + + # 2) Filing-mix survey over the window. + this_year = _dt.date.today().year + start_year = this_year - max(args.years - 1, 0) + date_range = f"{start_year}-01-01:{this_year}-12-31" + try: + df = company.get_filings(date=date_range).to_pandas() + except Exception as exc: + c.log(f"ERROR: could not survey filings for {date_range}: {exc}") + sys.exit(1) + + print(f"\n## Filing mix {start_year}-{this_year}") + if df is None or len(df) == 0: + print("- (no filings in this window; widen --years)") + return + + df = df.copy() + df["_date"] = df["filing_date"].astype(str) + + rows = [] + for form, grp in df.groupby("form"): + d = grp["_date"] + rows.append((str(form), len(grp), d.min(), d.max())) + rows.sort(key=lambda r: (r[1], r[3]), reverse=True) + + print("\n| Form | Count | Earliest | Latest |") + print("| :-- | --: | :-- | :-- |") + for form, n, lo, hi in rows: + print(f"| {form} | {n} | {lo} | {hi} |") + + # 3) Most recent filings. + n_recent = min(args.recent, len(df)) + print(f"\n## {n_recent} most recent filings") + recent = df.sort_values("_date", ascending=False).head(n_recent) + print("\n| Date | Form | Accession |") + print("| :-- | :-- | :-- |") + for _, row in recent.iterrows(): + print(f"| {row['_date']} | {row['form']} | {row.get('accession_number', '')} |") + + +if __name__ == "__main__": + main() diff --git a/skills/sec-edgar-skill/scripts/parse_financials.py b/skills/sec-edgar-skill/scripts/parse_financials.py index 9ec414f..be5c8f5 100644 --- a/skills/sec-edgar-skill/scripts/parse_financials.py +++ b/skills/sec-edgar-skill/scripts/parse_financials.py @@ -1,131 +1,324 @@ -"""Extract XBRL financial statements from an SEC report (annual or quarterly) to CSV. +"""Extract XBRL financial statements from one exact SEC report to CSV. -Resolves the report (10-K / 20-F / 40-F / 10-Q / 6-K) for the requested year and optional quarter, -parses its XBRL, and writes each statement to -``//____.csv``. Prints the -absolute path(s) to stdout. Run ``--help`` for all flags. +Select by accession, or by company and filing period. Period selection is strict: +the command never substitutes a more recent report. Ambiguous financial 6-Ks are +listed with their accessions so the caller can select one exactly. """ +from __future__ import annotations + import argparse import sys +from pathlib import Path import _common as c -# key -> (edgartools accessor on xbrl.statements, output filename suffix) STATEMENTS = { "income": ("income_statement", "income"), "balance": ("balance_sheet", "balance"), - "cashflow": ("cashflow_statement", "cashflow"), # note: no underscore in "cashflow" + "cashflow": ("cashflow_statement", "cashflow"), } +_FINANCIAL_TERMS = ( + "financial statement", + "financial results", + "interim results", + "quarterly results", + "half-year", + "half year", + "six month", + "nine month", + "earnings release", + "operating and financial review", +) -def main(): - p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) - p.add_argument("--ticker", required=True, help="Ticker, CIK, or company name.") - p.add_argument("--year", type=int, required=True, help="Calendar year of the report.") - p.add_argument( - "--quarter", - type=int, - choices=[1, 2, 3, 4], - help="Calendar quarter of the report (for 10-Q / 6-K).", - ) - p.add_argument( - "--form", - help="Form type, e.g. 10-K, 10-Q, 20-F, 40-F, 6-K. Defaults to 10-Q/6-K if quarter is specified, otherwise annual reports.", - ) - p.add_argument( - "--statement", - choices=["income", "balance", "cashflow", "all"], - default="all", - help="Which statement(s) to extract (default: all).", - ) - c.add_identity_arg(p) - c.add_cache_arg(p) - args = p.parse_args() - c.resolve_identity(args.identity) - company = c.resolve_company(args.ticker) +def _accession(filing) -> str: + return str(getattr(filing, "accession_no", "") or getattr(filing, "accession_number", "")) + + +def _available_statements(xbrl) -> list[str]: + available = [] + for key, (accessor, _) in STATEMENTS.items(): + method = getattr(xbrl.statements, accessor, None) + if method is None: + continue + try: + if method(): + available.append(key) + except Exception: + continue + return available + +def _xbrl_period(xbrl) -> str: + info = getattr(xbrl, "entity_info", None) or {} + fiscal_period = info.get("fiscal_period") + fiscal_year = info.get("fiscal_year") + period_end = getattr(xbrl, "period_of_report", None) or info.get("document_period_end_date") + label = " ".join(str(value) for value in (fiscal_period, fiscal_year) if value) + if period_end: + return f"XBRL period {label + ' ' if label else ''}ending {period_end}".strip() + return f"XBRL period {label}".strip() if label else "" + + +def _financial_evidence(filing) -> list[str]: + evidence = [] + if getattr(filing, "is_xbrl", False) or getattr(filing, "is_inline_xbrl", False): + evidence.append("XBRL metadata") + try: + for attachment in list(filing.attachments): + document = str(getattr(attachment, "document", "") or "") + description = str(getattr(attachment, "description", "") or "") + text = f"{document} {description}".lower() + matched = next((term for term in _FINANCIAL_TERMS if term in text), None) + if matched: + evidence.append(description or document or matched) + if "101.ins" in text or "inline xbrl" in text: + evidence.append("Inline XBRL attachment") + except Exception: + pass + return list(dict.fromkeys(evidence)) + + +def _log_candidates(title: str, rows: list[tuple[object, list[str], list[str]]]) -> None: + c.log(title) + c.log("Filed Form Accession Statements Evidence") + for filing, statements, evidence in rows: + c.log( + f"{getattr(filing, 'filing_date', '')!s:<10} " + f"{getattr(filing, 'form', '')!s:<8} " + f"{_accession(filing):<26} " + f"{','.join(statements) or '-':<16} " + f"{'; '.join(evidence) or '-'}" + ) + + +def _nearby_filings(company, forms: list[str]) -> None: + try: + nearby = list(company.get_filings(form=forms, amendments=False))[:5] + except Exception: + return + if not nearby: + return + c.log("Nearby original filings:") + for filing in nearby: + c.log(f" {filing.filing_date} {filing.form:<8} {_accession(filing)}") + + +def _select_company_filing(args: argparse.Namespace, company): + six_k_mode = False if args.form: forms = [args.form] + six_k_mode = args.form.upper() in {"6-K", "6-K/A"} + amendments = six_k_mode + filings = list( + company.get_filings( + form=forms, + year=args.year, + quarter=args.quarter, + amendments=amendments, + ) + ) elif args.quarter: - forms = ["10-Q", "6-K"] + forms = ["10-Q"] + filings = list( + company.get_filings( + form=forms, + year=args.year, + quarter=args.quarter, + amendments=False, + ) + ) + if not filings: + forms = ["6-K"] + six_k_mode = True + filings = list( + company.get_filings( + form=forms, + year=args.year, + quarter=args.quarter, + amendments=True, + ) + ) else: forms = ["10-K", "20-F", "40-F"] + filings = list(company.get_filings(form=forms, year=args.year, amendments=False)) - c.log( - f"Finding report ({'/'.join(forms)}) for {args.year}" - + (f" Q{args.quarter}" if args.quarter else "") - + " ..." - ) - - kwargs = {"form": forms, "year": args.year, "amendments": False} - if args.quarter: - kwargs["quarter"] = args.quarter - - filings = company.get_filings(**kwargs) - if len(filings) == 0: + if not filings: c.log( - f"No report found for {args.year}" + f"ERROR: no {'/'.join(forms)} filing found in filing year {args.year}" + (f" Q{args.quarter}" if args.quarter else "") - + f" of type {forms}; falling back to most recent." + + "." ) - kwargs_fallback = {"form": forms, "amendments": False} - filings = company.get_filings(**kwargs_fallback) - if len(filings) == 0: - c.log(f"ERROR: no reports found for {args.ticker} of type {forms}.") + _nearby_filings(company, forms) sys.exit(1) - filing = None - xbrl = None - filing_list = list(filings) - filing_list.sort(key=lambda f: f.filing_date, reverse=True) - - for f in filing_list: + parsed: list[tuple[object, object, list[str], list[str]]] = [] + likely_textual = [] + for filing in sorted(filings, key=lambda item: item.filing_date, reverse=True): + evidence = _financial_evidence(filing) if six_k_mode else [] try: - x = f.xbrl() - if x: - filing = f - xbrl = x - break + xbrl = filing.xbrl() + statements = _available_statements(xbrl) if xbrl else [] except Exception: - continue + xbrl = None + statements = [] + if xbrl and statements: + period = _xbrl_period(xbrl) + if period: + evidence.append(period) + parsed.append((filing, xbrl, statements, evidence)) + elif evidence: + likely_textual.append((filing, statements, evidence)) - if not filing or not xbrl: - c.log(f"ERROR: no filings with parsed XBRL data found for {args.ticker} in this period.") + if len(parsed) == 1: + filing, xbrl, statements, _ = parsed[0] + period = _xbrl_period(xbrl) + c.log( + f"Structured statements available: {', '.join(statements)}" + + (f"; {period}." if period else ".") + ) + return filing, xbrl + + if len(parsed) > 1: + _log_candidates( + "ERROR: multiple filings contain structured statements; rerun with --accession.", + [(filing, statements, evidence) for filing, _, statements, evidence in parsed], + ) sys.exit(1) - c.log( - f"Resolved {filing.form} filed {filing.filing_date} " - f"(accession {filing.accession_no}) containing XBRL data." + if six_k_mode and likely_textual: + _log_candidates( + "ERROR: no parseable XBRL statements found. These 6-Ks may contain textual " + "financial disclosures; fetch the relevant accession and inspect its attachments.", + likely_textual, + ) + else: + _log_candidates( + "ERROR: matching filing(s) were found, but none produced structured statements.", + [(filing, [], _financial_evidence(filing)) for filing in filings], + ) + sys.exit(1) + + +def _validate_args(parser: argparse.ArgumentParser, args: argparse.Namespace) -> None: + if args.accession: + conflicting = [ + flag + for flag, value in ( + ("--ticker", args.ticker), + ("--year", args.year), + ("--quarter", args.quarter), + ("--form", args.form), + ) + if value is not None + ] + if conflicting: + parser.error(f"--accession cannot be combined with {', '.join(conflicting)}") + elif not args.ticker or args.year is None: + parser.error("provide --accession, or provide both --ticker and --year") + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument( + "--accession", + help="Exact SEC accession. Cannot be combined with company/period selectors.", + ) + parser.add_argument("--ticker", help="Ticker, CIK, or company name.") + parser.add_argument("--year", type=int, help="Calendar filing year.") + parser.add_argument( + "--quarter", type=int, choices=[1, 2, 3, 4], help="Calendar filing quarter." ) + parser.add_argument( + "--form", + help="Form type. By default, infer annual forms or prefer 10-Q before scanning 6-K.", + ) + parser.add_argument( + "--statement", + choices=["income", "balance", "cashflow", "all"], + default="all", + help="Statement(s) to extract (default: all).", + ) + c.add_identity_arg(parser) + c.add_cache_arg(parser) + c.add_force_arg(parser) + args = parser.parse_args() + _validate_args(parser, args) + c.resolve_identity(args.identity) wanted = ["income", "balance", "cashflow"] if args.statement == "all" else [args.statement] - out_dir = c.company_dir(c.cache_root(args.cache_dir), company, ticker_hint=args.ticker) + + if args.accession: + filing = c.resolve_filing(args.accession) + company = c.company_for_filing(filing) + xbrl = None + else: + company = c.resolve_company(args.ticker) + try: + filing, xbrl = _select_company_filing(args, company) + except Exception as exc: + c.log(f"ERROR: could not select a financial filing: {exc}") + sys.exit(1) + + out_dir = c.company_dir( + c.cache_root(args.cache_dir), + company, + ticker_hint=args.ticker if not args.accession else None, + ) stem = c.filing_stem(filing) + paths = {key: out_dir / f"{stem}__{STATEMENTS[key][1]}.csv" for key in wanted} + + if not args.force and all( + path.is_file() and path.stat().st_size > 0 for path in paths.values() + ): + for path in paths.values(): + c.log(f"Using cached artifact: {path.resolve()}") + c.emit(path) + return + + if xbrl is None: + try: + xbrl = filing.xbrl() + except Exception as exc: + c.log(f"ERROR: accession {_accession(filing)} could not be parsed as XBRL: {exc}") + sys.exit(1) + if not xbrl: + c.log(f"ERROR: accession {_accession(filing)} contains no parsed XBRL data.") + sys.exit(1) - saved = [] + c.log(f"Resolved {filing.form} filed {filing.filing_date} (accession {_accession(filing)}).") + saved: list[Path] = [] + missing = [] for key in wanted: - accessor, suffix = STATEMENTS[key] + path = paths[key] + if not args.force and path.is_file() and path.stat().st_size > 0: + c.log(f"Using cached artifact: {path.resolve()}") + saved.append(path) + continue + accessor, _ = STATEMENTS[key] method = getattr(xbrl.statements, accessor, None) if method is None: - c.log(f"WARNING: {accessor} not available on this filing.") + missing.append(key) continue try: statement = method() if not statement: - c.log(f"WARNING: no data for {accessor}.") + missing.append(key) continue - df = statement.to_dataframe() - path = out_dir / f"{stem}__{suffix}.csv" - df.to_csv(path, index=True) + dataframe = statement.to_dataframe() + dataframe.to_csv(path, index=True) saved.append(path) - c.log(f" saved {path.name} ({len(df)} rows)") + c.log(f" saved {path.name} ({len(dataframe)} rows)") except Exception as exc: c.log(f"WARNING: {accessor} failed: {exc}") + missing.append(key) - if not saved: - c.log("ERROR: no statements could be extracted.") + if missing: + c.log(f"WARNING: unavailable statement(s): {', '.join(missing)}.") + if not saved or (args.statement != "all" and missing): + c.log("ERROR: the requested statement could not be extracted.") sys.exit(1) for path in saved: c.emit(path) diff --git a/skills/sec-edgar-skill/scripts/test_setup.py b/skills/sec-edgar-skill/scripts/test_setup.py index 05f992b..23e1ed0 100644 --- a/skills/sec-edgar-skill/scripts/test_setup.py +++ b/skills/sec-edgar-skill/scripts/test_setup.py @@ -1,81 +1,81 @@ -"""Diagnose the skill's environment: Python, dependencies, identity, cache dir. - -Run this first if anything misbehaves. Add ``--live`` to also make one real SEC -request and confirm end-to-end connectivity (and identity). Exit code is 0 when -everything required is in place, 1 otherwise. -""" - -import argparse -import os -import sys - -import _common as c - -REQUIRED = { - "edgar": "edgartools (SEC EDGAR client)", - "pandas": "pandas (dataframes / CSV export)", -} -OPTIONAL = { - "truststore": "truststore (system-trust TLS — only needed behind a corporate proxy)", -} - - -def main(): - p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) - c.add_identity_arg(p) - c.add_cache_arg(p) - p.add_argument( - "--live", action="store_true", help="Make one live SEC request to verify connectivity." - ) - args = p.parse_args() - - print("sec-edgar-skill — environment diagnostics") - print(f" Python: {sys.version.split()[0]} ({sys.executable})") - - ok = True - for mod, desc in REQUIRED.items(): - try: - m = __import__(mod) - print(f" [ok] {mod} {getattr(m, '__version__', '?')} — {desc}") - except Exception: - print(f" [MISSING] {mod} — {desc}") - ok = False - for mod, desc in OPTIONAL.items(): - try: - m = __import__(mod) - print(f" [ok] {mod} {getattr(m, '__version__', '?')} — {desc}") - except Exception: - print(f" [warn] {mod} not installed — {desc}") - - identity = args.identity or os.environ.get("EDGAR_IDENTITY") - if identity and "@" in identity: - print(f" [ok] SEC identity: {identity}") - else: - print( - " [MISSING] SEC identity — set $EDGAR_IDENTITY or pass --identity " - "(required for any fetch; missing it returns HTTP 403)." - ) - ok = False - - print(f" cache dir: {c.cache_root(args.cache_dir)}") - - if args.live: - if not ok: - print(" [skip] --live skipped: resolve the issues above first.") - else: - try: - c.resolve_identity(args.identity) - from edgar import Company - - n = len(Company("AAPL").get_filings(form="10-K", year=2023)) - print(f" [ok] live SEC request succeeded (AAPL 10-K 2023: {n} filing(s)).") - except Exception as exc: - print(f" [FAIL] live SEC request failed: {exc}") - ok = False - - print(f"STATUS: {'OK' if ok else 'ISSUES FOUND'}") - sys.exit(0 if ok else 1) - - -if __name__ == "__main__": - main() +"""Diagnose the skill's environment: Python, dependencies, identity, cache dir. + +Run this first if anything misbehaves. Add ``--live`` to also make one real SEC +request and confirm end-to-end connectivity (and identity). Exit code is 0 when +everything required is in place, 1 otherwise. +""" + +import argparse +import os +import sys + +import _common as c + +REQUIRED = { + "edgar": "edgartools (SEC EDGAR client)", + "pandas": "pandas (dataframes / CSV export)", +} +OPTIONAL = { + "truststore": "truststore (system-trust TLS — only needed behind a corporate proxy)", +} + + +def main(): + p = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + c.add_identity_arg(p) + c.add_cache_arg(p) + p.add_argument( + "--live", action="store_true", help="Make one live SEC request to verify connectivity." + ) + args = p.parse_args() + + print("sec-edgar-skill — environment diagnostics") + print(f" Python: {sys.version.split()[0]} ({sys.executable})") + + ok = True + for mod, desc in REQUIRED.items(): + try: + m = __import__(mod) + print(f" [ok] {mod} {getattr(m, '__version__', '?')} — {desc}") + except Exception: + print(f" [MISSING] {mod} — {desc}") + ok = False + for mod, desc in OPTIONAL.items(): + try: + m = __import__(mod) + print(f" [ok] {mod} {getattr(m, '__version__', '?')} — {desc}") + except Exception: + print(f" [warn] {mod} not installed — {desc}") + + identity = args.identity or os.environ.get("EDGAR_IDENTITY") + if identity and "@" in identity: + print(" [ok] SEC identity: configured") + else: + print( + " [MISSING] SEC identity — set $EDGAR_IDENTITY or pass --identity " + "(required for SEC requests; 13f.info convenience queries are exempt)." + ) + ok = False + + print(f" cache dir: {c.cache_root(args.cache_dir)}") + + if args.live: + if not ok: + print(" [skip] --live skipped: resolve the issues above first.") + else: + try: + c.resolve_identity(args.identity) + from edgar import Company + + n = len(Company("AAPL").get_filings(form="10-K", year=2023)) + print(f" [ok] live SEC request succeeded (AAPL 10-K 2023: {n} filing(s)).") + except Exception as exc: + print(f" [FAIL] live SEC request failed: {exc}") + ok = False + + print(f"STATUS: {'OK' if ok else 'ISSUES FOUND'}") + sys.exit(0 if ok else 1) + + +if __name__ == "__main__": + main()