Deterministic, Isolated, and Self-Healing Python Execution Runtime for AI Agents.
This open-source engine executes Python code generated by LLMs inside a disposable, isolated sandbox, validates it statically with the Python ast module, and recovers from runtime errors through a deterministic self-healing loop backed by local heuristics and plug-and-play LLM-guided patch generation (OpenAI, Anthropic Claude, DeepSeek, Ollama, Google Gemini).
| Feature | Vanilla subprocess |
Docker Container | CodeShield (This Engine) |
|---|---|---|---|
| Startup Overhead | ~5 – 10 ms | ~1,500 – 3,000 ms | Sub-second (~30–250 ms via uv) |
| Isolation Mechanism | None (Host Process) | Container Namespaces / cgroups | Ephemeral Virtualenv (tempfile + uv) |
| AST Security Gate | ❌ None | ❌ None | ✅ Static AST inspection (os.system, eval) |
| Silent Failure Detection | ❌ None | ❌ None | ✅ Regex scanning for empty DataFrames/NaNs |
| Self-Healing Loop | ❌ None | ❌ None | ✅ 3-Tier Traceback Diagnosis + LLM Patch (Any Provider) |
[LLM Generated Code]
│
▼
[AST Static Gate] ──(Syntax/Security Violation)──► [Validation Error Report]
│ (Passed)
▼
[uv Isolated Sandbox] ──(Runtime Error/Silent Failure)──► [Traceback Classifier]
│ │
│ (Clean Execution: exit 0) ▼
▼ [LLM Self-Healing (Any Provider) / Local Heuristic]
[Verified Output (JSON)] ◄──(AST Validated Patch)─────────────────┘
- Primary:
uv venvfor ultra-fast environment creation and package installation. - Fallback: native
python -m venv+pipwhenuvis unavailable, so the engine works out of the box on any machine. - Each execution lands in its own temporary workspace that is destroyed after use.
The engine parses every snippet with the standard ast module and rejects:
SyntaxErrors before execution.- Bare
except:/except Exception:/except BaseException:handlers. - Calls to dangerous parametrizable functions:
eval(),exec(),compile(). - Calls to system/subprocess primitives:
os.system(),subprocess.call(),subprocess.run(),subprocess.Popen().
Even when a process exits with 0, the engine flags suspicious output patterns such as:
empty DataFrameall NaNTracebackPipeline failedFatal Error
AST Validation ──► Sandbox Execution ──► Traceback Classification ──► Patch ──► Re-run
(3 attempts max)
- Local heuristic fallback: handles
NameError,ImportError,ModuleNotFoundErrorby injecting safe imports or placeholder definitions. - LLM-guided healing: when an LLM is configured (built-in Gemini Flash by default, or any custom provider via
patch_generator), it asks the model for a corrected version of the code, validates it with the AST gate, and re-executes the patched snippet.
# Install from PyPI
pip install codeshield-runtime
# Install with all extras (LLM + Dev tools)
pip install "codeshield-runtime[llm,dev]"
# Or clone for development
git clone https://github.com/AlgorithmicMind/codeshield.git
cd codeshield
# With uv (recommended)
uv venv
uv pip install -e ".[test,lint,llm,dev]"
# Or with pip
python -m venv .venv
.venv\Scripts\activate # Windows
pip install -e ".[test,lint,llm,dev]"from codeshield import SelfHealingEngine
# The sandbox is created and destroyed automatically on every ``run`` call.
engine = SelfHealingEngine(use_llm=False)
result, diagnosis = engine.run("print('hello world')")
print(result.stdout)
# Use a ``with`` block to reuse a single sandbox across multiple runs.
with SelfHealingEngine(use_llm=False) as reusable_engine:
first, _ = reusable_engine.run("print(1 + 1)")
second, _ = reusable_engine.run("print(2 + 2)")
print(first.stdout, second.stdout)CodeShield is not locked into a single LLM. Pass any Python callable as the patch_generator to use OpenAI, Anthropic Claude, DeepSeek, Ollama, LiteLLM or your own service:
from codeshield import SelfHealingEngine
def custom_llm_patcher(code: str, diagnosis) -> str:
# Compatible with any frontier provider: GPT-5.6, Claude Sonnet 5, DeepSeek V4, Ollama
response = client.chat.completions.create(
model="gpt-5.6-luna", # or "claude-sonnet-5", "deepseek-v4-flash"
messages=[
{
"role": "user",
"content": f"Fix this code:\n{code}\nError: {diagnosis.message}",
}
],
)
return response.choices[0].message.content
engine = SelfHealingEngine(patch_generator=custom_llm_patcher)
# Broken code -> AST gate -> sandbox execution -> traceback classification ->
# custom patch -> AST re-validation -> re-execution, all in a single call.
result, diagnosis = engine.run('print("Result: " + 42)')
print(result.stdout) # Result: 42For the built-in zero-config experience, create a .env file from .env.example:
GEMINI_API_KEY=your_key_here
GEMINI_MODEL=gemini-3.7-flash
from dotenv import load_dotenv
from codeshield import SelfHealingEngine
load_dotenv()
engine = SelfHealingEngine()
with engine:
result, diagnosis = engine.run('print("Result: " + 42)')
print(result.stdout) # Result: 42Run the included demo:
python demo.pyExecute any Python file directly from the terminal with the built-in CLI:
python -m codeshield run script.py
python -m codeshield run script.py --timeout 30
python -m codeshield run script.py --llm # try LLM self-healing if configured
python -m codeshield run script.py --no-llm # force local fallbackfrom codeshield import create_code_execution_tool
# Standard usage: built-in Gemini healing when configured, local heuristic otherwise
tools = [create_code_execution_tool()]
# Or bring your own model: the patcher is forwarded to the internal engine
tools = [create_code_execution_tool(patch_generator=custom_llm_patcher)]create_code_execution_tool() returns a ready-to-register execute_python_code(code: str) -> str function. It runs the provided Python in a self-healing sandbox and returns either the stdout or a structured error report with error_type and stderr.
The callable exposes real type hints and a Google-style docstring, so any SDK that builds a function schema from a plain Python callable can register it directly.
Everything is re-exported at the package root, so imports never need internal submodules:
from codeshield import (
ASTSecurityError,
CodeExecutionRequest,
ErrorDiagnosis,
ExecutionResult,
SandboxManager,
SelfHealingEngine,
SelfHealingError,
SubprocessRunner,
TracebackClassifier,
create_code_execution_tool,
validate_syntax_and_safety,
)ASTSecurityError subclasses SelfHealingError and is raised whenever the static AST gate blocks either the original snippet or a generated patch.
The examples/ folder contains ready-to-run recipes that have been executed and verified:
01_basic_sandboxing.py: isolated execution with timing measurements.02_security_gatekeeper.py: AST rejection of unsafe code.03_llm_healing_workflow.py: self-healing workflow with an LLM or local fallback.04_agent_tool_dropin.py: dual-phase agentic trace, printing agent thought, tool call, sandbox runtime and final answer for both a legitimate analytics round and a security-defense round where the AST gate blocks a shell escape.05_custom_llm_openai_compatible.py: model-agnostic, API-key-free self-healing with a custompatch_generator.
python examples/01_basic_sandboxing.py
python examples/02_security_gatekeeper.py
python examples/03_llm_healing_workflow.py
python examples/04_agent_tool_dropin.py
python examples/05_custom_llm_openai_compatible.pyThe suite currently has 54 tests with >83% code coverage on src/codeshield.
ruff check src tests examples
pytest tests -v --cov=src/codeshieldThis repository ships the core execution and healing engine. For production multi-tenant deployments, the enterprise extension adds:
- Multi-tenant orchestrator with queue-based job scheduling.
- PostgreSQL state persistence for execution history, audit trails and replay.
- Automated billing and token governance (cost caps per tenant, per-execution budgets).
- Prometheus/Grafana observability, RBAC, and signed artifact provenance.
- SLA-backed support and custom agentic architecture consulting.
Want the production-grade version or a tailored integration for your platform?
We offer enterprise licensing, dedicated onboarding and custom agentic-architecture consulting.
This project is licensed under the MIT License. See LICENSE for details.