Skip to main content

ABOUT ARCLUX

ARCLUX is an open-source platform that reads code the way an operating system reads hardware β€” and turns it into a living map of the repository. Point it at any codebase and it parses the source, computes how everything connects, and answers questions like β€œwhat breaks if I change this file?” or β€œthis file is never imported β€” where was it supposed to be wired?” β€” in seconds, without you reading 800 files first. ARCLUX is built in two layers:
  • The intelligence layer β€” the flagship capability: static analysis, dependency/call graphs, impact tracing, 20 automated code-health detectors, 14 framework convention rules, and a security pipeline. This is what works end-to-end today and is verified against real repos (vscode, react, vite, laravel, django, flask).
  • The platform layer β€” the runtime underneath: a kernel with a signal bus, process manager, job scheduler, service manager, storage, networking, notifications, and orchestration. This is the foundation ARCLUX grows on β€” codebase intelligence is the first application of the platform, not the last.
Everything in this file is real and working. No marketing promises β€” every claim is backed by code and verified runs.

The map

The intelligence layer β€” what you can do with it

The 20 detectors

Automated code-health checks, each an independent small file that is trivial to extend:
  • circularDependency β€” import cycles, with the full cycle path
  • unusedExports / unusedFiles β€” code that nothing consumes
  • orphanFiles β€” files nothing imports, classified: dead (leftover, delete it) vs unwired (should be connected) vs ambiguous
  • orphanIntegration β€” for unwired files, where they should be imported: the folder’s barrel index, or the shared importer of same-kind siblings (confidence + score + evidence, derived from real patterns β€” never guessed)
  • largeModules / duplicateModules / sharedModules / indexFiles β€” structural smell detection
  • layerViolation β€” imports that cross architecture layers
  • deadCode / ambiguousSymbolResolution / missingExports
  • componentConvention / featureStructure / repositoryPattern / routeConvention / storyConvention / testConvention / entryPoints

Remote sources & security boundaries

arclux analyze https://github.com/org/repo clones, analyzes, and cleans up. Source adapters route any input β€” GitHub, GitLab (https/ssh/SCP-style), archive files, local paths (~ expanded) β€” through the right boundary check:
  • SSRF guard β€” remote URLs are refused before anything else if they point at private networks, loopback, link-local, or cloud metadata endpoints (169.254.169.254). Public hosts β€” GitHub, GitLab, Bitbucket, any web server β€” are always allowed.
  • Source boundary β€” local paths are checked against allowed/denied roots with symlink containment.
  • Evidence boundary β€” doctor/security output is redacted (tokens, keys, passwords, AWS credentials, private keys, connection strings) and per-check capped.
  • Analysis boundary β€” hard caps on files/bytes/modules so no single run can exhaust the host.

What ARCLUX understands

  • Languages parsed today: TypeScript/TSX, JavaScript, Python, Go, Java, PHP, Ruby, Rust, C++, C#, Bash, C, Dart, Elixir, Kotlin, Lua, Objective-C, OCaml, Scala, Solidity, Swift, Vue, Zig, Elm, ReScript β€” via TypeScript Compiler API + web-tree-sitter (25 grammar-backed, 2 compiler-API-backed, plus manifest parsers for package.json, go.mod, Cargo.toml, Gemfile, composer.json, csproj, gradle, pom.xml, requirements.txt)
  • Frameworks with convention rules: Next.js, NestJS, Express, Vite, Electron, React, Laravel
  • Graphs: dependency (imports/exports/folders) and call graph (which function calls which, across files) + folder graph

The ARCLUX DSL β€” scripting the analysis

arclux script <file.arclux> runs a tiny scripting language purpose-built for codebase intelligence. Scripts read like instructions, not API calls:
Every capability the engine exposes is bound into the DSL β€” analyze, doctor, check, graph, callgraph, impact, search, security, diff, archdiff, plus helpers (len, sum, filter, sort, exists, keys, values, env, cwd, extensions, checkids). The language grows automatically: registering a new parser or detector expands extensions() / checkids() with zero DSL changes (verified live β€” the 5 new parsers from PR #528 grew the binding surface from 9 to 19 extensions on their own). The browser playground at /script runs the same DSL server-side β€” including an audit mode that streams doctor + security + attack-surface findings as a terminal theater and replays them as breathing halos on the 3D dependency graph.

How it works (the 10-second version)

Each stage is an independent package (parser, graph, impact, detectors, rules, engine, security). Add a new parser, detector, or rule without touching the rest. The single entry point is analyzeRepository in the engine pipeline β€” nothing calls individual steps from outside.

The platform layer

Beneath the intelligence layer sits a real runtime, not scaffolding:
  • kernel β€” a signal bus every subsystem emits and subscribes through (Kernel)
  • runtime β€” ProcessManager spawns and supervises child processes (RuntimeManager)
  • scheduler β€” JobScheduler + JobQueue for async work
  • services β€” ServiceManager manages service lifecycle and dependencies
  • storage β€” ArtifactStore, CacheManager, RecoveryManager (crash-safe writes), SnapshotManager
  • networking β€” ConnectionManager, PortManager, ServiceEndpoint discovery files
  • notifications β€” NotificationManager fans events out to channels
  • orchestration β€” PlatformOrchestrator assembles it all
The daemon, watcher, incremental indexer, shell session, and workspace layers sit on top of these. Packages like observation, web-intake, and package-manager mark the direction the platform is heading β€” not yet wired, but the seams are already there.

What ARCLUX does NOT do yet (honest)

  • Analysis history is persisted per-run (JSON-record store wired into the daemon), but there’s no query layer over it yet β€” packages/db has schema + stores, higher-level queries aren’t built
  • Per-file incremental re-indexing: the incremental engine is built, but buildIndex still does a full rebuild per change β€” per-file wiring is deferred
  • The platform’s runtime layers (scheduler/services/storage/observation/web-intake) are built but not all wired to consumers
  • Some exotic tree-sitter grammars shipped in tree-sitter-wasms are stale (elm was ABI 12 β€” vendored fix; ReScript’s wasm predates its modern import syntax) β€” see packages/parser/wasms/

Fair Evaluation Protocol β€” how to judge ARCLUX (or any architecture)

Don’t evaluate architecture from screenshots. Screenshot β†’ feeling β†’ ranking is not analysis. The fair ladder is:
  1. Repository β€” look at structure and scope, not renders.
  2. Source β€” read the actual implementation.
  3. Graph β€” check dependencies, impact, coupling (arclux graph, arclux impact, arclux doctor do exactly this).
  4. Build/Test β€” verify claims executably (build scripts, tsc, detectors, checks β€” not words).
  5. Runtime β€” observe real system behavior (server, snapshots, persistence β€” not mockups).
  6. Conclusion β€” only conclude what the evidence supports.
And keep four words apart β€” an honest engineer knows exactly which is which:

Where to go next

  • QUICKSTART.md β€” fast-path workflow cheat sheet
  • CONTEXT.md β€” stack, architecture, current state at a glance
  • docs site β€” full, searchable documentation
  • PROGRES.md (+ progres/) β€” live status: what works, what’s a stub, known bugs