Cluster size distribution test — below p_c, does the finite cluster size distribution follow a power law with exponent…
executedProtocol p-b-habitat-percolation-ecology-cluster-exponent · desktop
Researchers & contributors: Start here → · Launch Sprint guide · 12 easy-win issues →
The Universal Science Discovery Repository is open infrastructure for the scientific unknown. A git-native, schema-validated catalog of 1,409 open research problems, 1,275 falsifiable hypotheses, and 1,124 cross-domain mathematical bridges — all connected in a reproducible 3,861-node knowledge graph. Built so researchers can find high-leverage open questions fast, institutions can spot interdisciplinary opportunities, and funders can see exactly where the frontier is moving.
Find entries: full-text search (titles and claims), filter by domain chip, then explore the interactive knowledge graph. Press / anywhere to focus search after you scroll to it.
Updated by python scripts/update_dashboard_stats.py --apply (catalog globs + docs/knowledge_graph.json meta). Wave milestones stay in ROADMAP.md; bounded PR-sized content batches and validation reminders are in docs/PATH_TO_SUCCESS.md.
build_graph.py
Same path as CONTRIBUTING.md and docs/ONBOARDING.md, laid out as steps. Clone the repo so file links open directly in your editor.
Plain-language map of the tools you will touch — no scientific claims, just workflow.
| Tool | What it does | You use it to… |
|---|---|---|
| Hub (this page) | Browse catalog, search, graph, hygiene list | Find a task; read only |
| Git + GitHub | Version control | Edit YAML; open PR |
validate_schemas + repo_smoke |
Local/CI checks | Confirm YAML valid before PR |
| build-graph bot PR | Auto refresh graph/API after merge | Merge when it appears |
| GitHub Pages | Hosted hub | Share link; green banner = site current |
Deeper walkthrough: CONTRIBUTING.md · docs/DEV_DASHBOARD.md
The hub browses and routes you to tasks; your PR edits YAML in a local clone. Stream A walkthrough: HAPPY_PATH_FIRST_RECORDS.md ↗
Get a local copy so you can edit catalog files and run validation scripts.
Choose one entry path — all end in a pull request.
Use the hygiene table, Stream A templates, or the issue’s file path.
From the repo root:
python scripts/validate_schemas.py
python -m pytest tests/repo_smoke
After your PR merges, watch for an optional graph rebuild bot PR and merge it if it appears (refreshes the knowledge graph JSON).
Vision, motivation, and why this infrastructure matters for science.
What lives here, what doesn't, and the rigor expected of every entry.
The evidence bar, claims discipline, and output conventions. Read before proposing any scientific claims.
Ethics policy, what data may be committed, and community norms.
Stream A: validate setup → add a u-… unknown → add a h-… hypothesis → open PR.
USDR maps what connects; Crosscheck tests it. Each protocol links to a repro bundle — run in your browser, open in Colab, or clone for verification. Crosscheck manifesto
Protocol p-b-habitat-percolation-ecology-cluster-exponent · desktop
Protocol p-b-habitat-percolation-ecology-fss · desktop
Protocol p-b-ising-social-dynamics-ewi · desktop
Protocol p-b-percolation-epidemiology-fss · desktop
Real-time filter across all unknowns, bridges, hypotheses, and phenomena loaded from the knowledge graph. Press / to focus.
Read-only list built from the same catalog scan as the knowledge graph (not a scientific ranking).
Fixing stale IDs and missing links keeps the graph truthful for everyone.
Regenerate api/v1/orphan_xref_panel.json with
python scripts/export_orphan_xref_panel.py
after catalog edits — see docs/DEV_DASHBOARD.md.
Broken cross-reference means a YAML file points at an ID that does not exist (typo, renamed record, or missing target). The graph cannot draw that link until the ID is fixed or the reference is removed.
Quick fix (three steps):
related_* ID in that YAML file so it matches a real catalog ID.This table shows at most 100 rows from the export — not every xref issue in the repo.
Loading contribution targets…
Six curated thematic lenses — shortcut entry points into representative bridges in the catalog, not exhaustive domain coverage.
Maintainer playbook (regenerate stats, domain pages, breakthrough cards): docs/DEV_DASHBOARD.md.
Foundational scientists whose work seeds cross-domain bridges — including underappreciated contributions ripe for rediscovery.
AC power · Wireless energy · Bladeless turbine · Earth-ionosphere resonance
Unified electromagnetism · Statistical mechanics · Maxwell's demon
S = k ln W · Arrow of time · Statistical foundations
Symmetry → conservation laws · Abstract algebra · Gauge invariance
Information entropy · Channel capacity · Boolean circuits
Turing machine · Halting problem · Morphogenesis reaction-diffusion
QED path integrals · Feynman diagrams · Quantum computing pioneer
Transposable elements · Genome plasticity · Epigenetic regulation
Natural selection · Common descent · Sexual selection · Earthworm geology
Photo 51 · DNA X-ray crystallography · Virus structure · Carbon microstructure
Stored-program computer · Game theory · Quantum formalism · Self-reproducing automata
Special & general relativity · Photoelectric effect · Brownian motion · Bose–Einstein condensate
First algorithm · General-purpose computing vision · Analytical Engine programming
World-reshaping breakthroughs stalled by cross-domain knowledge gaps. Cards below are generated from the breakthrough-gaps catalog (same YAML CI validates). Click a card to open the source YAML on GitHub; Alt-click (Option-click on macOS) jumps to Catalog search with a prefilled query.
Direct air capture (DAC) at ~$100/tonne CO₂ (from current $400–1,000/tonne) would make it economically viable to remove atmospheric CO₂ at the gigatonne scale required to meaningfully address climate change. At $100/ton…
The COVID-19 mRNA vaccines (BioNTech/Pfizer, Moderna, 2020–2021) proved the mRNA platform at TRL 9: design-to-manufacture in months, high efficacy, large-scale production. The platform is now being extended to personali…
High-bandwidth brain-computer interfaces (BCIs) — capable of reading and writing neural activity at the resolution of individual neurons across large brain areas — would enable restoration of motor function in paralysis…
Nuclear fusion — combining hydrogen isotopes to release energy via E=mc² — offers effectively unlimited, low-carbon electricity with no long-lived nuclear waste and inherent safety (no chain reaction, self-extinguishing…
Detecting any cancer type at stage I from a blood draw — using circulating tumor DNA (ctDNA), cell-free methylation signatures, exosomes, or protein biomarkers — with less than 1% false positive rate and greater than 90…
Scalable, low-cost water splitting using earth-abundant catalysts operating near room temperature would make green hydrogen competitive with natural gas reformation, enabling carbon-neutral fuels, long-duration grid sto…
In December 2022, the National Ignition Facility achieved Q_fusion > 1 for the first time — the fusion reaction released more energy than the laser energy deposited in the fuel capsule. However, accounting for wall-plug…
Natural photosynthesis captures ~1-2% of incident solar energy as chemical energy in biomass — orders of magnitude below the theoretical maximum efficiency (~11% for oxygenic photosynthesis, limited by thermodynamics of…
Scientific literature grows at approximately 4% per year, doubling every 17 years. As of 2024, PubMed alone indexes over 37 million articles; the total corpus of peer-reviewed science across all fields exceeds 100 milli…
Alzheimer's disease (AD) affects 55 million people worldwide and is the leading cause of dementia. After 25 years and >100 failed clinical trials, the amyloid hypothesis (Hardy & Higgins 1992) received partial vindicati…
A fault-tolerant quantum computer with ~1,000 error-corrected logical qubits would break RSA-2048 encryption, simulate quantum chemistry at pharmaceutical-design accuracy (protein folding, drug binding), and solve optim…
Current quantum computers have 1,000-10,000 physical qubits but physical error rates of 0.1-1% per gate, far above the threshold needed for useful computation. Quantum error correction using surface codes requires appro…
Aging is the largest single risk factor for cancer, heart disease, neurodegeneration, and metabolic disease, yet its mechanism is debated. The leading candidate theories include: epigenetic entropy (information theory o…
Antimicrobial resistance (AMR) is projected to cause 10 million deaths per year by 2050 (O'Neill Review 2016), exceeding cancer mortality. The WHO classifies AMR as one of the top 10 global public health threats. The la…
AlphaFold2 and ESMFold solve the forward problem — predicting 3D structure from amino acid sequence — with near-experimental accuracy. The inverse problem remains largely unsolved: designing a novel sequence that will f…
A material that superconducts at room temperature and ambient pressure would eliminate resistive losses in electrical grids (~5–10% of generated power), enable compact MRI and fusion magnets without cryogenic infrastruc…
The hard problem of consciousness — why and how physical processes in the brain give rise to subjective experience (qualia) — has no agreed scientific framework. The easy problems (explaining cognitive functions like at…
No technology can simultaneously record all ~86 billion neurons in a human brain at single-neuron resolution during natural behavior. Current state-of-the-art (Neuropixels probes) records roughly 10,000 neurons simultan…
Current operational global climate models run at 25-100 km horizontal resolution. At this resolution, key processes governing regional climate — mesoscale convective systems, cumulus convection, cloud microphysics, orog…
The human brain performs general-purpose intelligence at approximately 20 watts. Running a 70-billion parameter large language model requires ~70 watts per token generated at the GPU level, and a full inference data cen…
Borrelia burgdorferi, the Lyme disease spirochete, causes persistent debilitating symptoms in 10–20% of patients even after standard antibiotic treatment. PTLDS (sometimes called chronic Lyme disease) involves fatigue,…
Soil contains approximately 10^9 microorganisms per gram, representing more than 10,000 species per sample interacting through metabolic exchange, competition, and mutualism in networks too complex to model from first p…
Materials that can autonomously reconfigure their macroscopic shape, stiffness, and function in response to external commands do not exist beyond proof-of-concept demonstrations at millimeter scale. Shape memory alloys…
Approximately 170 trillion plastic particles are estimated to be in the ocean, with an additional 8-10 million tons entering annually. Macro-plastic removal by systems such as The Ocean Cleanup is technically feasible b…
Three scripts that continuously mine the knowledge graph — surfacing gaps, proposing novel cross-domain connections, and flagging low-quality entries for human review.
Domain pairs with high unknown density and no existing bridge — the most fertile candidates for the next cross-domain discovery.
View proposals ↗Unknowns with no hypothesis or bridge edge — the highest-impact contribution opportunities in the graph today.
View targets ↗Automated quality checks across all 401 catalog entries — errors, warnings, and improvement opportunities surfaced automatically.
View audit ↗Pattern-matching engine that detects which mathematical framework (phase transitions, information theory, scaling laws) best connects two domains, then drafts bridge YAMLs for expert review.
How it works →Papers cited across multiple bridges — the most cross-domain influential works in the scientific literature. Shannon, Turing, Fisher, and their equivalents.
View index →
No authentication. No rate limits. Served directly from GitHub Pages.
Base URL: https://kr8zysho3.github.io/Universal-Science-Discovery/api/v1/
GET /api/v1/meta.json
Catalog statistics, counts, and metadata
~1 KB
GET /api/v1/bridges.json
All 1124 cross-domain bridges with translation tables and references
~213 KB
GET /api/v1/unknowns.json
All 1409 open unknowns with domain, status, and file path
~451 KB
GET /api/v1/hypotheses.json
All 1275 hypotheses with priority, impact score, and evidence links
~114 KB
GET /api/v1/breakthrough_gaps.json
24 breakthrough-gap summaries (TRL, potential, blocking / required-bridge counts)
~8 KB
GET /api/v1/domains.json
Per-domain summary statistics (unknowns and bridge counts)
~2 KB
GET /api/v1/graph.json
Full knowledge graph — 3861 nodes, 4522 edges (built from catalog YAML; mirrors docs/knowledge_graph.json)
~578 KB
GET /api/v1/bridge_proposals.json
AI co-pilot bridge proposals ranked by novelty score
~6 KB
GET /api/v1/orphan_xref_panel.json
Capped list of missing xref targets and disconnected unknowns for the hub panel (regenerate via export script)
~40 KB
Force-directed graph of the full catalog — bridges, unknowns, hypotheses, and phenomena (counts update live when the JSON loads). Hover a node to inspect it. Click a node for full details. Drag to reposition. Scroll or pinch to zoom.
Development is divided into independent areas. Find one that fits your skills, comment on an open issue to claim it, and open a draft PR early. Full details: WORKSTREAMS.md.
status:needs-owner. Comment "I'm taking this." A maintainer assigns it to you.feat/<area>/<slug> from main. Open a draft PR immediately to signal what you are working on.main is branch-protected — no direct pushes.Phase 0 — Foundation is complete (governance, schemas, CI, catalog seed, graph, hub). Below are calendar- and community-dependent milestones; development and catalog growth continue in parallel.
Ring = completed Phase 1 items only. In-progress items still appear in the checklist below.
This hub does not edit the catalog — it links to the folders and guides where you add YAML in your clone. Start with HAPPY_PATH_FIRST_RECORDS.md for your first unknown + hypothesis PR.
Step-by-step: validate setup → add u-… unknown → add h-… hypothesis → open PR.
Research gaps tracked as u-… YAML files under unknowns-catalog/.
Testable proposals as h-… YAML with evidence links and falsification criteria.
Explicit connections between fields studying the same phenomenon — the anti-tunnel-vision layer. Schema: b-… YAML.
YAML schemas for validation and PR/issue templates for structured contributions.
Rules of the road, AI use policy, and ingest integration notes.
If you use Cursor or other agents — rules and agent-specific policy.
Vision, roadmap, system architecture, and outreach framing.
Default landing for new visitors; overview of all major areas.
North-star vision through 2035, phase milestones, and guiding principles.
Discovery Core, Human Layer, AI Layer, and the Integration & Data Layer.
Accurate language for external conversations and contributor recruiting.
Governance, legal framework, conduct, and the quality bar that makes contributions credible.
What the repo may host; third-party attribution and IP policy.
Decision-making structure, ethics framework, and integrity requirements.
CI gates, review lanes, definition of done — the anti-sloppiness playbook.
Review expectations, issue labels, and working group norms.
Branch protection, CI gates, and the daily operating rhythm.
From "this file" to "what policy it enforces" — audit trails and the full onboarding path.
Phase A/B metadata plans, ingest envelope schema, and example data.
Doc discipline: at each milestone or feature merge, update README,
CHANGELOG (Unreleased), relevant docs/, and this hub if links change. ·
Pull requests ·
Open issues ·
CHANGELOG.md