jshookmcp is a free, open source data extraction & web scraping project written in TypeScript and released under AGPL-3.0. It has 2,020 GitHub stars, 459 forks and 0 open issues, and was last pushed 20 hours ago. On this registry it ranks #69 of 122 tracked projects in Data Extraction & Web Scraping, with 5 head-to-head comparisons available.

What is jshookmcp?

jshook is an AGPL-3.0, strict-TypeScript reverse-engineering workspace distributed as a Model Context Protocol (MCP) server under the package name @jshookmcp/jshook, built for AI agents performing JavaScript and WASM reverse engineering, network analysis, deobfuscation and process instrumentation.

What it is

jshook is a search-first, profile-aware reverse-engineering workspace for AI agents. It organises its surface into 36 self-discovered domains and 733 tools, exposed through three profiles: "search", which loads roughly 3K tokens of tool metadata ranked by BM25 and hybrid vector search; "workflow", which runs composite scripts; and "full", which exposes all 733 tools at around 40K tokens. Browser-side work runs over CDP against Chromium and Camoufox; deeper analysis reaches into Frida, Ghidra, IDA, wabt, Binaryen and Burp Suite. Plugins hot-reload, workflows are declarative, and auto-discovery lets the tool surface grow without a redeploy. It lives in the MCP ecosystem, on the Node.js/TypeScript toolchain, and in this registry under Data & Analytics / Data Extraction & Web Scraping.

The concrete problem it solves is the trade-off most MCP servers for JavaScript analysis impose: expose a handful of hand-rolled tools, or wrap a single browser engine. jshook replaces both with a full stack, so hooking a page, capturing network traffic, deobfuscating a bundle, disassembling WASM and instrumenting a process sit behind one server rather than a set of regex calls dressed up as tools. It also solves the token cost of that breadth: an agent starts on the small "search" profile and escalates through "search" → "workflow" → "full" only as the task demands, instead of drowning the model in schemas from the first turn.

Key capabilities

  • Three tool profiles over 36 self-discovered domains: "search" (~3K tokens, BM25 plus hybrid vector ranking), "workflow" (composite scripts) and "full" (all 733 tools, ~40K tokens).
  • Browser automation with Chromium and Camoufox over CDP, including attach to existing targets, anti-detection presets, an explicit-input CAPTCHA solver, and popup, download, permission and protocol interceptors.
  • Network interception covering HTTP/1.1 and HTTP/2 frame building, a MITM proxy with an on-demand auto-generated CA, WebSocket capture, GraphQL introspection helpers and a Burp Suite bridge.
  • JavaScript hooks and analysis: LLM-powered deobfuscation, crypto routine detection, AST comprehension and transforms, source-map reconstruction, and script/scriptlet extraction and replay.
  • WASM reverse engineering through wabt (wasm2wat, wasm-decompile, wasm-objdump, wasm2c), Binaryen's wasm-opt, a pure-TypeScript section parser, function-level binary diff, obfuscation detection, and function- and basic-block-level instrumentation.
  • Process and memory forensics: native FFI scanning, cross-reference graphs, hardware breakpoints, PE introspection, live process attach, and memory read/write with region guards.
  • Runtime recovery and session isolation: streamable HTTP sessions restore activated domains, browser attach state and coverage state after reconnects, while per-client browser-side state stays isolated.

Who uses it and how

  • Security analysts and reverse engineers taking apart obfuscated front-end bundles, using LLM-powered deobfuscation, crypto routine detection and source-map reconstruction in a single session.
  • Analysts attached to an existing, possibly broken page through CDP, relying on runtime recovery to restore attach and coverage state after a reconnect.
  • Investigators pulling apart compiled WASM payloads, listing imports and exports, recovering the name section, extracting strings by section, and diffing function-level binary changes.
  • Binary and malware analysts working on PE files and live processes with native FFI scanning, hardware breakpoints, and the Frida, Ghidra and IDA bridges.
  • Multi-agent setups: because per-client browser-side state is isolated, two agents can work concurrently without trampling each other's CDP sessions.

Getting started

jshook is distributed as the package @jshookmcp/jshook, and requires Node.js 22.22.2+ or 24.15+, with pnpm 10.x used for the build. Installation, configuration and the full tool reference are covered in the README's Getting Started and Configuration sections and on the documentation site at https://vmoranv.github.io/jshookmcp/.

How it compares

No comparable products are named in the available facts, so jshook stands alone in this registry, with no direct paid equivalents listed for contrast. The README positions it against the general run of MCP servers for JS analysis, which it characterises as exposing a handful of hand-rolled tools or wrapping a single browser engine.

When to use it — and when not to

A self-hoster must run a Node.js 22.22.2+ or 24.15+ service, build with pnpm 10.x, and operate a streamable HTTP MCP endpoint; the deeper capabilities additionally assume external toolchains such as Frida, Ghidra, IDA, wabt and Binaryen are available. The AGPL-3.0 licence imposes copyleft obligations that make it a poor fit for teams needing to embed it in closed-source products, and the "full" profile's 733 tools at roughly 40K tokens of metadata is a real context cost for smaller models. Teams wanting a managed, hosted service rather than a server they run themselves should look elsewhere.

project readme (upstream, from github) — read inline

@jshookmcp/jshook

License: AGPLv3 Node.js 22.22.2+ TypeScript MCP pnpm

A search-first, profile-aware reverse-engineering workspace for AI agents.

Hook the page, capture the network, deobfuscate the bundle, disassemble the WASM, instrument the process — and let one MCP server keep the whole attack surface in reach without drowning the model in schemas.

English · 中文

Stars Forks Latest Release License

Node.js 22.22.2+ TypeScript strict MCP current

What's different · Capabilities · Use cases · Highlights · Transport · Registry · Architecture · Build

Documentation · Getting Started · Configuration · Tool Reference

Sponsored by Swiftproxy — Premium Residential Proxies for Web Automation · 10% off code: PROXY90


What makes jshook different

Most MCP servers for JS analysis expose a handful of hand-rolled tools or wrap a single browser engine. jshook is closer to an operating system for front-end reverse engineering — 36 self-discovered domains, a search-first meta-tool that keeps token cost under control, and runtime recovery that survives broken pages and dropped sessions:

  • Search-first, profile-aware. The search profile loads about 3K tokens of tool metadata; the full profile exposes all 733 tools at around 40K tokens. Agents move between them as the task grows — search → workflow → full — instead of drowning in schemas from the first turn.
  • Runtime recovery and session isolation. Streamable HTTP sessions restore activated domains, browser attach state, and coverage state after reconnects; per-client browser-side state stays isolated so two agents cannot trample each other's CDP sessions.
  • Full-stack browser automation. Chromium and Camoufox via CDP with anti-detection, an explicit-input CAPTCHA solver (no built-in page/feature probing), a self-signed HTTPS interception CA on demand, and HTTP/2 frame building.
  • Real reverse engineering, not string searches. WASM disassembly via wabt (wasm2wat / wasm-decompile / wasm-objdump), Frida/Ghidra/IDA bridges, native FFI scanning, hardware breakpoints, PE introspection, GraphQL/Burp Suite proxy bridges, and AST transforms — not a single regex call wrapped as a tool.
  • Dynamic extensibility. Hot-reload plugins, declarative workflows, and auto-discovery keep the server growing without a redeploy.

Capability overview

A scan of what's in the box. Each row links to the detailed Capability overview below.

Area Highlights
Tool profiles search (~3K tokens, BM25 + hybrid vector ranking) · workflow (composite scripts) · full (all 733 tools)
Browser automation Chromium and Camoufox · CDP attach to existing targets · anti-detection presets · explicit-input CAPTCHA solver · popup, download, permission, and protocol interceptors
Network interception HTTP/1.1 + HTTP/2 frame building · MITM proxy with auto-generated CA · WebSocket capture · GraphQL introspection helpers · Burp Suite bridge
JS hooks and analysis LLM-powered deobfuscation · crypto routine detection · AST comprehension · source-map reconstruction · script/scriptlet extraction and replay
WASM reverse engineering wabt disassembly / C transpilation (wasm2wat / wasm-decompile / wasm-objdump / wasm2c) · Binaryen wasm-opt · pure-TS section parser · import/export listing · section-grouped string extraction with name-section recovery · function-level binary diff · obfuscation detection · function- and basic-block-level instrumentation
Process and memory forensics Native FFI scanning · cross-reference graphs · hardware breakpoints · PE introspection · live process attach · memory read/write with region guards
Binary instrumentation Frida bridge · Ghidra and IDA bridges · syscall hooking · TLS keylog and session tooling · Mojo IPC analysis
Android and APK analysis APK static triage · manifest dump and query · apktool decode / build · jadx decompilation and code search · DEX scanning · native library listing · APK signing · unidbg emulation · runtime DEX dump
Native runtime Native emulator for foreign-architecture samples · platform introspection · Mojo IPC · Dart Inspector · ADB bridge for on-device traffic
Encoding and transform URL/Base64/Hex/JWT/Protobuf encoders · AST transforms · streaming decode pipelines
Coordination Background task queue with progress, cancellation, and async modes · multi-agent coordination · coverage reports
Schema-first meta tools describe_tool · call_tool with argument validation · coverage_report · search_tools
Pluggable extension registry Hot-reload plugins · declarative workflows · auto-discovered domains

Use cases

Scenario What you do Domains involved
Skim a minified bundle search_tools → deobfuscate → search_in_scripts → understand_code core
Reverse a CAPTCHA challenge Drive a Camoufox page → screenshot → solve with explicit input → replay browser, canvas
Capture and replay an OAuth flow proxy_start (auto CA) → network_get_requests → graphql_introspect → graphql_replay proxy, network, graphql
Reverse a WASM crypto routine wasm_dump → wasm_disassemble → generate_hooks → memory_breakpoint wasm, binary-instrument, memory
Triage a suspicious APK apk_static_triage → apk_manifest_dump → jadx_decompile → dex_scan_file binary-instrument
Recover a dropped browser session Reconnect Streamable HTTP → restore activated domains and browser state browser, coordination
Audit a Node process for credentials process_list → memory_scan_filtered → binary_strings_extract process, memory, binary-instrument
Build a custom workflow list_extension_workflows → run_extension_workflow workflow, extension-registry
Hook a function in a live process frida_spawn → frida_attach_interceptor → frida_run_script → frida_enumerate_functions binary-instrument

Quick start

No global install needed — add to your MCP client config and you're ready.

Claude Desktop / Cursor (claude_desktop_config.json):

{
  "mcpServers": {
    "jshook": {
      "command": "npx",
      "args": ["-y", "@jshookmcp/jshook@latest"],
      "env": {
        "MCP_TOOL_PROFILE": "search",
        "npm_config_omit": "optional"
      }
    }
  }
}

(Windows: use npx.cmd absolute path if npx is not found.)

This lightweight configuration skips optional ONNX, Z3, Binaryen, Camoufox, and Playwright packages. Remove npm_config_omit when those full-profile runtimes are required.

Share one daemon across multiple agents

The default stdio configuration starts one full jshook process per MCP host. To share the embedding model, browser runtime, and caches, start one local Streamable HTTP daemon:

pnpm build
pnpm daemon

Vector search defaults to off for per-client stdio processes and on (lazy-loaded) for the shared HTTP daemon. Set SEARCH_VECTOR_ENABLED=false when lexical search is sufficient.

Then point every MCP client at http://127.0.0.1:3000/mcp using its HTTP/URL server configuration. Each client receives its own MCP session and response route while heavyweight runtime resources remain in one process. Keep the default loopback bind; set MCP_AUTH_TOKEN before exposing the endpoint beyond localhost.

Promote a profile as the task grows

{
  "env": {
    "MCP_TOOL_PROFILE": "search"     // start here, ~3K tokens of metadata
  }
}

Switch MCP_TOOL_PROFILE to workflow once you start chaining composite scripts, or to full when you need every tool. coverage_report shows the active set on demand.


Highlights

  • Profile ladder. Start in search (~3K tokens of metadata); promote to workflow when chaining composite scripts; escalate to full only when every tool is actually needed. coverage_report shows what's active on demand.
  • Meta tools. describe_tool returns the JSON Schema; call_tool validates arguments before invocation; every tool ships with readOnlyHint / destructiveHint / idempotentHint / openWorldHint.
  • Browser automation. Chromium and Camoufox via CDP, attach to existing targets, anti-detection presets, popup/download/permission interceptors, explicit-input CAPTCHA solver, JS/CSS injection at three document phases, persisted coverage across reconnects.
  • Network interception. Auto-generated HTTPS interception CA, HTTP/1.1 + HTTP/2 frame building, WebSocket capture, GraphQL helpers, Burp Suite bridge — all on the same MCP tool surface.
  • Reverse engineering. wabt WASM disassembly (wasm2wat / wasm-decompile / wasm-objdump) and Binaryen wasm-opt, Frida/Ghidra/IDA bridges, hardware breakpoints, native FFI scanning, PE introspection, syscall hooking, AST transforms, source-map reconstruction.
  • Session recovery. Streamable HTTP transport restores activated domains, browser attach state, and coverage state after reconnects; browser-side state is isolated per client.
  • Plugins and workflows. Drop a directory, get a domain. Write a YAML pipeline, run it as one tool. The registry self-discovers.

Transport and deployment

The server supports two transports out of the box.

Transport When to use Notes
stdio Default for Claude Desktop, Cursor, and other single-host clients One full process per MCP host; lightweight profile recommended
Streamable HTTP Multiple agents sharing the embedding model, browser runtime, and caches Loopback bind by default; set MCP_AUTH_TOKEN before exposing externally

Both transports expose the same tool surface. coverage_report shows which domains are activated in each session — long-running sessions restore browser attach state, coverage state, and tool activations across reconnects.

For production deployments see the Security and Production guide.


Recent runtime notes

  • HTTP transport now multiplexes independent MCP sessions and restores runtime state after reconnects.
  • proxy_start auto-generates a local HTTPS interception CA when needed.
  • Browser CAPTCHA solving is now explicit-input driven: pass taskKind, siteKey, imageBase64, callbackName, and responseSelector as needed. Built-in widget/page signature probing is intentionally not used.

Registry snapshot

The built-in surface below is generated from the runtime registry and checked in CI.

  • Package version: 0.3.5
  • Built-in tools: 733
  • Domains: adb-bridge, binary-instrument, browser, canvas, coordination, core, cross-domain, dart-inspector, debugger, encoding, exploit-dev, extension-registry, graphql, instrumentation, maintenance, memory, mojo-ipc, native-bridge, native-emulator, network, platform, process, protocol-analysis, proxy, session, sourcemap, streaming, syscall-hook, tasks, tls-inspector, trace, transform, v8-inspector, wasm, webgpu, workflow
  • Note: this snapshot is generated from the runtime registry; do not edit the counts by hand.

View the complete Tool Reference ↗


Architecture

  • Runtime registry — domains auto-discovered via manifest.ts; add a domain by creating one file.
  • Lazy initialization — handlers instantiated on first call, not at startup.
  • BM25 + vector search — search_tools meta-tool with hybrid ranking and adaptive weights.
  • MCP ToolAnnotations — every tool carries readOnlyHint / destructiveHint / idempotentHint / openWorldHint.
  • Profile ladder — search (~3K tokens) → workflow (composite scripts) → full (all 733 tools).
  • Transport symmetry — stdio and Streamable HTTP expose the same surface; sessions are isolated per client.

See the Architecture guide and Configuration reference for the canonical details.


Build from source

Requirements: Node.js 22.22.2+, pnpm 10.x.

pnpm install
pnpm build
pnpm start           # run the built server from dist/
pnpm dev             # run from source under tsx watch
pnpm check           # drift guards (metadata + openapi + domain + event contracts) + lint + format check + typecheck + unit tests
pnpm test            # Vitest unit suites
pnpm test:e2e        # end-to-end browser/tooling suites
pnpm daemon          # run the Streamable HTTP daemon after build

Native helpers are bundled via pnpm build; on first run the server may download optional runtimes (ONNX, Z3, Binaryen, Camoufox, Playwright) depending on the profile.


Project stats

Star History Chart

Activity


License

AGPLv3.

Frequently asked questions

Is jshookmcp free to use?

jshookmcp is open source under the AGPL-3.0 licence. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does jshookmcp do?

js hook toolkit that all you need

What is jshookmcp written in?

jshookmcp is primarily written in TypeScript. Its source is publicly available at https://github.com/vmoranv/jshookmcp, and it has 2,020 GitHub stars.