diff --git a/docs/CONTRIBUTING.md b/docs/CONTRIBUTING.md
new file mode 100644
index 0000000000000000000000000000000000000000..de43641f6737cdcb95128d4f7e98bacc145adb8c
--- /dev/null
+++ b/docs/CONTRIBUTING.md
@@ -0,0 +1,571 @@
+# How to contribute
+
+We would love to accept your patches and contributions to this project. This
+document includes:
+
+- **[Before you begin](#before-you-begin):** Essential steps to take before
+ becoming a Gemini CLI contributor.
+- **[Code contribution process](#code-contribution-process):** How to contribute
+ code to Gemini CLI.
+- **[Development setup and workflow](#development-setup-and-workflow):** How to
+ set up your development environment and workflow.
+- **[Documentation contribution process](#documentation-contribution-process):**
+ How to contribute documentation to Gemini CLI.
+
+We're looking forward to seeing your contributions!
+
+## Before you begin
+
+### Sign our Contributor License Agreement
+
+Contributions to this project must be accompanied by a
+[Contributor License Agreement](https://cla.developers.google.com/about) (CLA).
+You (or your employer) retain the copyright to your contribution; this simply
+gives us permission to use and redistribute your contributions as part of the
+project.
+
+If you or your current employer have already signed the Google CLA (even if it
+was for a different project), you probably don't need to do it again.
+
+Visit to see your current agreements or to
+sign a new one.
+
+### Review our Community Guidelines
+
+This project follows
+[Google's Open Source Community Guidelines](https://opensource.google/conduct/).
+
+## Code contribution process
+
+### Get started
+
+The process for contributing code is as follows:
+
+1. **Find an issue** that you want to work on. If an issue is tagged as
+ `πMaintainers only`, this means it is reserved for project maintainers. We
+ will not accept pull requests related to these issues. In the near future,
+ we will explicitly mark issues looking for contributions using the
+ `help-wanted` label. If you believe an issue is a good candidate for
+ community contribution, please leave a comment on the issue. A maintainer
+ will review it and apply the `help-wanted` label if appropriate. Only
+ maintainers should attempt to add the `help-wanted` label to an issue.
+2. **Fork the repository** and create a new branch.
+3. **Make your changes** in the `packages/` directory.
+4. **Ensure all checks pass** by running `npm run preflight`.
+5. **Open a pull request** with your changes.
+
+### Code reviews
+
+All submissions, including submissions by project members, require review. We
+use [GitHub pull requests](https://docs.github.com/articles/about-pull-requests)
+for this purpose.
+
+To assist with the review process, we provide an automated review tool that
+helps detect common anti-patterns, testing issues, and other best practices that
+are easy to miss.
+
+#### Using the automated review tool
+
+You can run the review tool in two ways:
+
+1. **Using the helper script (Recommended):** We provide a script that
+ automatically handles checking out the PR into a separate worktree,
+ installing dependencies, building the project, and launching the review
+ tool.
+
+ ```bash
+ ./scripts/review.sh [model]
+ ```
+
+ **Warning:** If you run `scripts/review.sh`, you must have first verified
+ that the code for the PR being reviewed is safe to run and does not contain
+ data exfiltration attacks.
+
+ **Authors are strongly encouraged to run this script on their own PRs**
+ immediately after creation. This allows you to catch and fix simple issues
+ locally before a maintainer performs a full review.
+
+ **Note on Models:** By default, the script uses the latest Pro model
+ (`gemini-3.1-pro-preview`). If you do not have enough Pro quota, you can run
+ it with the latest Flash model instead:
+ `./scripts/review.sh gemini-3-flash-preview`.
+
+2. **Manually from within Gemini CLI:** If you already have the PR checked out
+ and built, you can run the tool directly from the CLI prompt:
+
+ ```text
+ /review-frontend
+ ```
+
+Replace `` with your pull request number. Reviewers should use this
+tool to augment, not replace, their manual review process.
+
+### Self-assigning and unassigning issues
+
+To assign an issue to yourself, simply add a comment with the text `/assign`. To
+unassign yourself from an issue, add a comment with the text `/unassign`.
+
+The comment must contain only that text and nothing else. These commands will
+assign or unassign the issue as requested, provided the conditions are met
+(e.g., an issue must be unassigned to be assigned).
+
+Please note that you can have a maximum of 3 issues assigned to you at any given
+time and that only
+[issues labeled "help wanted"](https://github.com/google-gemini/gemini-cli/issues?q=is%3Aissue%20state%3Aopen%20label%3A%22help%20wanted%22)
+may be self-assigned.
+
+### Pull request guidelines
+
+To help us review and merge your PRs quickly, please follow these guidelines.
+PRs that do not meet these standards may be closed.
+
+#### 1. Link to an existing issue
+
+All PRs should be linked to an existing issue in our tracker. This ensures that
+every change has been discussed and is aligned with the project's goals before
+any code is written.
+
+- **For bug fixes:** The PR should be linked to the bug report issue.
+- **For features:** The PR should be linked to the feature request or proposal
+ issue that has been approved by a maintainer.
+
+If an issue for your change doesn't exist, we will automatically close your PR
+along with a comment reminding you to associate the PR with an issue. The ideal
+workflow starts with an issue that has been reviewed and approved by a
+maintainer. Please **open the issue first** and wait for feedback before you
+start coding.
+
+#### 2. Keep it small and focused
+
+We favor small, atomic PRs that address a single issue or add a single,
+self-contained feature.
+
+- **Do:** Create a PR that fixes one specific bug or adds one specific feature.
+- **Don't:** Bundle multiple unrelated changes (e.g., a bug fix, a new feature,
+ and a refactor) into a single PR.
+
+Large changes should be broken down into a series of smaller, logical PRs that
+can be reviewed and merged independently.
+
+#### 3. Use draft PRs for work in progress
+
+If you'd like to get early feedback on your work, please use GitHub's **Draft
+Pull Request** feature. This signals to the maintainers that the PR is not yet
+ready for a formal review but is open for discussion and initial feedback.
+
+#### 4. Ensure all checks pass
+
+Before submitting your PR, ensure that all automated checks are passing by
+running `npm run preflight`. This command runs all tests, linting, and other
+style checks.
+
+#### 5. Update documentation
+
+If your PR introduces a user-facing change (e.g., a new command, a modified
+flag, or a change in behavior), you must also update the relevant documentation
+in the `/docs` directory.
+
+See more about writing documentation:
+[Documentation contribution process](#documentation-contribution-process).
+
+#### 6. Write clear commit messages and a good PR description
+
+Your PR should have a clear, descriptive title and a detailed description of the
+changes. Follow the [Conventional Commits](https://www.conventionalcommits.org/)
+standard for your commit messages.
+
+- **Good PR title:** `feat(cli): Add --json flag to 'config get' command`
+- **Bad PR title:** `Made some changes`
+
+In the PR description, explain the "why" behind your changes and link to the
+relevant issue (e.g., `Fixes #123`).
+
+### Forking
+
+If you are forking the repository you will be able to run the Build, Test and
+Integration test workflows. However in order to make the integration tests run
+you'll need to add a
+[GitHub Repository Secret](https://docs.github.com/en/actions/security-for-github-actions/security-guides/using-secrets-in-github-actions#creating-secrets-for-a-repository)
+with a value of `GEMINI_API_KEY` and set that to a valid API key that you have
+available. Your key and secret are private to your repo; no one without access
+can see your key and you cannot see any secrets related to this repo.
+
+Additionally you will need to click on the `Actions` tab and enable workflows
+for your repository, you'll find it's the large blue button in the center of the
+screen.
+
+### Development setup and workflow
+
+This section guides contributors on how to build, modify, and understand the
+development setup of this project.
+
+### Setting up the development environment
+
+**Prerequisites:**
+
+1. **Node.js**:
+ - **Development:** Please use Node.js `~20.19.0`. This specific version is
+ required due to an upstream development dependency issue. You can use a
+ tool like [nvm](https://github.com/nvm-sh/nvm) to manage Node.js versions.
+ - **Production:** For running the CLI in a production environment, any
+ version of Node.js `>=20` is acceptable.
+2. **Git**
+
+### Build process
+
+To clone the repository:
+
+```bash
+git clone https://github.com/google-gemini/gemini-cli.git # Or your fork's URL
+cd gemini-cli
+```
+
+To install dependencies defined in `package.json` as well as root dependencies:
+
+```bash
+npm install
+```
+
+To build the entire project (all packages):
+
+```bash
+npm run build
+```
+
+This command typically compiles TypeScript to JavaScript, bundles assets, and
+prepares the packages for execution. Refer to `scripts/build.js` and
+`package.json` scripts for more details on what happens during the build.
+
+### Enabling sandboxing
+
+[Sandboxing](#sandboxing) is highly recommended and requires, at a minimum,
+setting `GEMINI_SANDBOX=true` in your `~/.env` and ensuring a sandboxing
+provider (e.g. `macOS Seatbelt`, `docker`, or `podman`) is available. See
+[Sandboxing](#sandboxing) for details.
+
+To build both the `gemini` CLI utility and the sandbox container, run
+`build:all` from the root directory:
+
+```bash
+npm run build:all
+```
+
+To skip building the sandbox container, you can use `npm run build` instead.
+
+### Running the CLI
+
+To start the Gemini CLI from the source code (after building), run the following
+command from the root directory:
+
+```bash
+npm start
+```
+
+If you'd like to run the source build outside of the gemini-cli folder, you can
+utilize `npm link path/to/gemini-cli/packages/cli` (see:
+[docs](https://docs.npmjs.com/cli/v9/commands/npm-link)) or
+`alias gemini="node path/to/gemini-cli/packages/cli"` to run with `gemini`
+
+### Running tests
+
+This project contains two types of tests: unit tests and integration tests.
+
+#### Unit tests
+
+To execute the unit test suite for the project:
+
+```bash
+npm run test
+```
+
+This will run tests located in the `packages/core` and `packages/cli`
+directories. Ensure tests pass before submitting any changes. For a more
+comprehensive check, it is recommended to run `npm run preflight`.
+
+#### Integration tests
+
+The integration tests are designed to validate the end-to-end functionality of
+the Gemini CLI. They are not run as part of the default `npm run test` command.
+
+To run the integration tests, use the following command:
+
+```bash
+npm run test:e2e
+```
+
+For more detailed information on the integration testing framework, please see
+the
+[Integration Tests documentation](https://geminicli.com/docs/integration-tests).
+
+### Linting and preflight checks
+
+To ensure code quality and formatting consistency, run the preflight check:
+
+```bash
+npm run preflight
+```
+
+This command will run ESLint, Prettier, all tests, and other checks as defined
+in the project's `package.json`.
+
+_ProTip_
+
+after cloning create a git precommit hook file to ensure your commits are always
+clean.
+
+```bash
+echo "
+# Run npm build and check for errors
+if ! npm run preflight; then
+ echo "npm build failed. Commit aborted."
+ exit 1
+fi
+" > .git/hooks/pre-commit && chmod +x .git/hooks/pre-commit
+```
+
+#### Formatting
+
+To separately format the code in this project, run the following command from
+the root directory:
+
+```bash
+npm run format
+```
+
+This command uses Prettier to format the code according to the project's style
+guidelines.
+
+#### Linting
+
+To separately lint the code in this project, run the following command from the
+root directory:
+
+```bash
+npm run lint
+```
+
+### Coding conventions
+
+- Please adhere to the coding style, patterns, and conventions used throughout
+ the existing codebase.
+- Consult
+ [GEMINI.md](https://github.com/google-gemini/gemini-cli/blob/main/GEMINI.md)
+ (typically found in the project root) for specific instructions related to
+ AI-assisted development, including conventions for React, comments, and Git
+ usage.
+- **Imports:** Pay special attention to import paths. The project uses ESLint to
+ enforce restrictions on relative imports between packages.
+
+### Debugging
+
+#### VS Code
+
+0. Run the CLI to interactively debug in VS Code with `F5`
+1. Start the CLI in debug mode from the root directory:
+ ```bash
+ npm run debug
+ ```
+ This command runs `node --inspect-brk dist/gemini.js` within the
+ `packages/cli` directory, pausing execution until a debugger attaches. You
+ can then open `chrome://inspect` in your Chrome browser to connect to the
+ debugger.
+2. In VS Code, use the "Attach" launch configuration (found in
+ `.vscode/launch.json`).
+
+Alternatively, you can use the "Launch Program" configuration in VS Code if you
+prefer to launch the currently open file directly, but 'F5' is generally
+recommended.
+
+To hit a breakpoint inside the sandbox container run:
+
+```bash
+DEBUG=1 gemini
+```
+
+**Note:** If you have `DEBUG=true` in a project's `.env` file, it won't affect
+gemini-cli due to automatic exclusion. Use `.gemini/.env` files for gemini-cli
+specific debug settings.
+
+### React DevTools
+
+To debug the CLI's React-based UI, you can use React DevTools.
+
+1. **Start the Gemini CLI in development mode:**
+
+ ```bash
+ DEV=true npm start
+ ```
+
+2. **Install and run React DevTools version 6 (which matches the CLI's
+ `react-devtools-core`):**
+
+ You can either install it globally:
+
+ ```bash
+ npm install -g react-devtools@6
+ react-devtools
+ ```
+
+ Or run it directly using npx:
+
+ ```bash
+ npx react-devtools@6
+ ```
+
+ Your running CLI application should then connect to React DevTools.
+ 
+
+### Sandboxing
+
+#### macOS Seatbelt
+
+On macOS, `gemini` uses Seatbelt (`sandbox-exec`) under a `permissive-open`
+profile (see `packages/cli/src/utils/sandbox-macos-permissive-open.sb`) that
+denies operations by default, confining writes to the project folder while
+allowing broad file reads and outbound network traffic ("open") by default. You
+can switch to a `strict-open` profile (see
+`packages/cli/src/utils/sandbox-macos-strict-open.sb`) that restricts both reads
+and writes to the working directory while allowing outbound network traffic by
+setting `SEATBELT_PROFILE=strict-open` in your environment or `.env` file.
+Available built-in profiles are `permissive-{open,proxied}`,
+`restrictive-{open,proxied}`, and `strict-{open,proxied}` (see below for proxied
+networking). You can also switch to a custom profile
+`SEATBELT_PROFILE=` if you also create a file
+`.gemini/sandbox-macos-.sb` under your project settings directory
+`.gemini`.
+
+#### Container-based sandboxing (all platforms)
+
+For stronger container-based sandboxing on macOS or other platforms, you can set
+`GEMINI_SANDBOX=true|docker|podman|` in your environment or `.env`
+file. The specified command (or if `true` then either `docker` or `podman`) must
+be installed on the host machine. Once enabled, `npm run build:all` will build a
+minimal container ("sandbox") image and `npm start` will launch inside a fresh
+instance of that container. The first build can take 20-30s (mostly due to
+downloading of the base image) but after that both build and start overhead
+should be minimal. Default builds (`npm run build`) will not rebuild the
+sandbox.
+
+Container-based sandboxing mounts the project directory (and system temp
+directory) with read-write access and is started/stopped/removed automatically
+as you start/stop Gemini CLI. Files created within the sandbox should be
+automatically mapped to your user/group on host machine. You can easily specify
+additional mounts, ports, or environment variables by setting
+`SANDBOX_{MOUNTS,PORTS,ENV}` as needed. You can also fully customize the sandbox
+for your projects by creating the files `.gemini/sandbox.Dockerfile` and/or
+`.gemini/sandbox.bashrc` under your project settings directory (`.gemini`) and
+running `gemini` with `BUILD_SANDBOX=1` to trigger building of your custom
+sandbox.
+
+#### Proxied networking
+
+All sandboxing methods, including macOS Seatbelt using `*-proxied` profiles,
+support restricting outbound network traffic through a custom proxy server that
+can be specified as `GEMINI_SANDBOX_PROXY_COMMAND=`, where ``
+must start a proxy server that listens on `:::8877` for relevant requests. See
+`docs/examples/proxy-script.md` for a minimal proxy that only allows `HTTPS`
+connections to `example.com:443` (e.g. `curl https://example.com`) and declines
+all other requests. The proxy is started and stopped automatically alongside the
+sandbox.
+
+### Manual publish
+
+We publish an artifact for each commit to our internal registry. But if you need
+to manually cut a local build, then run the following commands:
+
+```
+npm run clean
+npm install
+npm run auth
+npm run prerelease:dev
+npm publish --workspaces
+```
+
+## Documentation contribution process
+
+Our documentation must be kept up-to-date with our code contributions. We want
+our documentation to be clear, concise, and helpful to our users. We value:
+
+- **Clarity:** Use simple and direct language. Avoid jargon where possible.
+- **Accuracy:** Ensure all information is correct and up-to-date.
+- **Completeness:** Cover all aspects of a feature or topic.
+- **Examples:** Provide practical examples to help users understand how to use
+ Gemini CLI.
+
+### Getting started
+
+The process for contributing to the documentation is similar to contributing
+code.
+
+1. **Fork the repository** and create a new branch.
+2. **Make your changes** in the `/docs` directory.
+3. **Preview your changes locally** in Markdown rendering.
+4. **Lint and format your changes.** Our preflight check includes linting and
+ formatting for documentation files.
+ ```bash
+ npm run preflight
+ ```
+5. **Open a pull request** with your changes.
+
+### Documentation structure
+
+Our documentation is organized using
+[sidebar.json](https://github.com/google-gemini/gemini-cli/blob/main/docs/sidebar.json)
+as the table of contents. When adding new documentation:
+
+1. Create your markdown file **in the appropriate directory** under `/docs`.
+2. Add an entry to `sidebar.json` in the relevant section.
+3. Ensure all internal links use relative paths and point to existing files.
+
+### Style guide
+
+We follow the
+[Google Developer Documentation Style Guide](https://developers.google.com/style).
+Please refer to it for guidance on writing style, tone, and formatting.
+
+#### Key style points
+
+- Use sentence case for headings.
+- Write in second person ("you") when addressing the reader.
+- Use present tense.
+- Keep paragraphs short and focused.
+- Use code blocks with appropriate language tags for syntax highlighting.
+- Include practical examples whenever possible.
+
+### Linting and formatting
+
+We use `prettier` to enforce a consistent style across our documentation. The
+`npm run preflight` command will check for any linting issues.
+
+You can also run the linter and formatter separately:
+
+- `npm run lint` - Check for linting issues
+- `npm run format` - Auto-format markdown files
+- `npm run lint:fix` - Auto-fix linting issues where possible
+
+Please make sure your contributions are free of linting errors before submitting
+a pull request.
+
+### Before you submit
+
+Before submitting your documentation pull request, please:
+
+1. Run `npm run preflight` to ensure all checks pass.
+2. Review your changes for clarity and accuracy.
+3. Check that all links work correctly.
+4. Ensure any code examples are tested and functional.
+5. Sign the
+ [Contributor License Agreement (CLA)](https://cla.developers.google.com/) if
+ you haven't already.
+
+### Need help?
+
+If you have questions about contributing documentation:
+
+- Check our [FAQ](https://geminicli.com/docs/resources/faq).
+- Review existing documentation for examples.
+- Open [an issue](https://github.com/google-gemini/gemini-cli/issues) to discuss
+ your proposed changes.
+- Reach out to the maintainers.
+
+We appreciate your contributions to making Gemini CLI documentation better!
diff --git a/docs/behavioral-evals.md b/docs/behavioral-evals.md
new file mode 100644
index 0000000000000000000000000000000000000000..823be20bf9372dce9ecb28aaf4739de6b701b409
--- /dev/null
+++ b/docs/behavioral-evals.md
@@ -0,0 +1,185 @@
+# Behavioral Evaluations & EDK Guide
+
+This guide introduces the **Eval Development Kit (EDK)** and details how to
+write, validate, run, and report on **behavioral evaluations** in the Gemini CLI
+codebase.
+
+---
+
+## Overview
+
+Behavioral evaluations are automated tests designed to assert on the
+**behavior** of the Gemini CLI agent (e.g., verifying which tools are called,
+checking call ordering, or avoiding destructive commands) rather than checking
+the final prose output.
+
+Evaluating agent behavior is critical because:
+
+1. Model responses are non-deterministic, making exact prose matching highly
+ fragile.
+2. We must ensure the model utilizes the most efficient tools (e.g., batching
+ files via `read_many_files` instead of sequential `read_file` calls).
+3. We must enforce safety boundaries (e.g., preventing execution of raw shell
+ commands when safe alternatives exist).
+
+All behavioral evaluations are stored under the `evals/` directory.
+
+---
+
+## EDK Developer Commands
+
+The EDK provides CLI tools under `scripts/` to help contributors audit, check,
+and monitor evals.
+
+### 1. `npm run eval:inventory`
+
+Scans all eval files under `evals/`, statically parses them, and provides a
+structured overview of what exists in the repository.
+
+- **Usage:**
+ ```bash
+ npm run eval:inventory
+ ```
+- **JSON Output:** For CI integration or inventory indexing, generate a
+ machine-readable JSON report:
+ ```bash
+ npm run eval:inventory -- --json
+ ```
+- **Custom Root:** Run against another directory or repository:
+ ```bash
+ npm run eval:inventory -- --root /path/to/other/repo
+ ```
+
+---
+
+### 2. `npm run eval:validate`
+
+A lint-like checker that validates eval source files against standard structural
+guidelines and best practices.
+
+- **Usage:**
+ ```bash
+ npm run eval:validate
+ ```
+- **Custom Scopes:** Validate a specific file:
+ ```bash
+ npm run eval:validate -- evals/my-test.eval.ts
+ ```
+
+#### Validation Rules & Severities
+
+| Rule ID | Severity | Description |
+| :------------------- | :---------- | :--------------------------------------------------------------------------------------------------------------------- |
+| `file-naming` | **Error** | File must match `*.eval.ts` or `*.eval.tsx` naming conventions. |
+| `valid-policy` | **Error** | Policy must be one of `ALWAYS_PASSES`, `USUALLY_PASSES`, or `USUALLY_FAILS`. |
+| `suite-metadata` | **Error** | Both `suiteName` and `suiteType` must be present as static string literals. |
+| `prompt-presence` | **Error** | Every eval case must have a non-empty `prompt` string. |
+| `case-name-static` | **Error** | The case name must be a static string literal, not computed dynamically. |
+| `invalid-tool-refs` | **Error** | All tools referenced in assertions must match known built-in or legacy tools. |
+| `positive-assertion` | **Error** | Evaluation cases must assert on at least one tool call (e.g., check `waitForToolCall` has been invoked). |
+| `workspace-setup` | **Error** | Workspace behaviors (like file-system edits/reads) must set up a `files` object. |
+| `new-evals-policy` | **Warning** | New evals must not use `ALWAYS_PASSES` policy initially (they should be promoted after nightly data proves stability). |
+
+Warnings (`new-evals-policy`) will be logged with `β ` and will **not** cause
+the CLI process to exit with status `1`. Errors (`β`) will block CI builds and
+return exit status `1`.
+
+---
+
+### 3. `npm run eval:report`
+
+Aggregates local vitest `report.json` artifacts, maps them against inventory
+policies, and summarizes the pass rates per model.
+
+- **Usage:**
+ ```bash
+ npm run eval:report
+ ```
+ By default, it scans `evals/logs/` recursively for `report.json` files.
+- **Specifying Directory:**
+ ```bash
+ npm run eval:report -- /path/to/logs
+ ```
+- **JSON Output:**
+ ```bash
+ npm run eval:report -- --json
+ ```
+
+---
+
+## Contributor Workflow
+
+When writing a new behavioral evaluation, adhere to this workflow to ensure
+high-quality, non-flaky test runs.
+
+### Step-by-Step Guide
+
+1. **Identify the Target Behavior**: Determine which tool calls need
+ verification (e.g., `web_fetch` must be called).
+2. **Author the Eval File**: Create your file under `evals/.eval.ts`
+ naming it properly.
+3. **Configure Workspace Files**: If the eval reads or edits files, define them
+ inside the `files` metadata field.
+4. **Assert Behavior, Not Prose**: Ensure the `assert` block checks tool
+ interactions using `rig.waitForToolCall` or similar. Do not check final
+ prose.
+5. **Run Locally**:
+ ```bash
+ RUN_EVALS=true npx vitest run evals/my-test.eval.ts
+ ```
+6. **Deflake**: Run your eval at least 3 times locally to verify it does not
+ fail due to model variance.
+7. **Run Validation**: Run `npm run eval:validate` to ensure no linting errors
+ are present.
+
+### Acceptance Criteria Checklist
+
+- [ ] **Naming**: File ends with `.eval.ts` or `.eval.tsx`.
+- [ ] **Policy**: New evals start as `USUALLY_PASSES`.
+- [ ] **Metadata**: Static `suiteName` and `suiteType` (e.g. `'behavioral'`) are
+ specified.
+- [ ] **Assertions**: Uses `rig.waitForToolCall` or asserts tool arguments
+ explicitly.
+- [ ] **Clean workspace**: Does not write to files outside `rig.testDir`.
+
+### Common Anti-Patterns to Avoid
+
+- **Restricting core tools**: Never override `settings.tools.core` to limit
+ tools. Evals must run against the default toolset.
+- **Checking model prose**: Avoid `expect(result).toContain('something')` since
+ model wording is non-deterministic.
+- **Integration-only testing**: Evals that only write files without checking
+ realistic model prompts are integration tests and belong under
+ `integration-tests/`.
+
+---
+
+## CI & Dashboard Integration
+
+You can easily automate behavioral evaluations or compile dashboard data using
+EDK's JSON reporters.
+
+### CI Validation Block
+
+Add a step in your PR checks or GitHub workflows to automatically lint new evals
+and block pull requests containing validation errors:
+
+```yaml
+- name: Run Eval Validator
+ run: npm run eval:validate
+```
+
+### Publishing to a Dashboard
+
+To record nightly performance metrics across multiple models:
+
+1. Configure your workflow to run evaluations with the JSON reporter:
+ ```bash
+ cross-env GEMINI_MODEL=gemini-2.5-pro npx vitest run --config evals/vitest.config.ts --reporter=json --outputFile="evals/logs/eval-logs-gemini-2.5-pro/report.json"
+ ```
+2. Aggregate all test runs using the reporting tool:
+ ```bash
+ npm run eval:report -- evals/logs --json > aggregated_report.json
+ ```
+3. Upload `aggregated_report.json` to your dashboard storage backend to
+ visualize pass rates over time.
diff --git a/docs/index.md b/docs/index.md
new file mode 100644
index 0000000000000000000000000000000000000000..d5fefce47c41816a5b7e8757789404e7e9a11a17
--- /dev/null
+++ b/docs/index.md
@@ -0,0 +1,139 @@
+# Gemini CLI documentation
+
+Gemini CLI brings the power of Gemini models directly into your terminal. Use it
+to understand code, automate tasks, and build workflows with your local project
+context.
+
+## Install
+
+```bash
+npm install -g @google/gemini-cli
+```
+
+## Get started
+
+Jump in to Gemini CLI.
+
+- **[Quickstart](./get-started/index.md):** Your first session with Gemini CLI.
+- **[Installation](./get-started/installation.mdx):** How to install Gemini CLI
+ on your system.
+- **[Authentication](./get-started/authentication.mdx):** Setup instructions for
+ personal and enterprise accounts.
+- **[CLI cheatsheet](./cli/cli-reference.md):** A quick reference for common
+ commands and options.
+- **[Gemini 3 on Gemini CLI](./get-started/gemini-3.md):** Learn about Gemini 3
+ support in Gemini CLI.
+
+## Use Gemini CLI
+
+User-focused guides and tutorials for daily development workflows.
+
+- **[File management](./cli/tutorials/file-management.md):** How to work with
+ local files and directories.
+- **[Get started with Agent skills](./cli/tutorials/skills-getting-started.md):**
+ Getting started with specialized expertise.
+- **[Manage context and memory](./cli/tutorials/memory-management.md):**
+ Managing persistent instructions and facts.
+- **[Execute shell commands](./cli/tutorials/shell-commands.md):** Executing
+ system commands safely.
+- **[Manage sessions and history](./cli/tutorials/session-management.md):**
+ Resuming, managing, and rewinding conversations.
+- **[Plan tasks with todos](./cli/tutorials/task-planning.md):** Using todos for
+ complex workflows.
+- **[Web search and fetch](./cli/tutorials/web-tools.md):** Searching and
+ fetching content from the web.
+- **[Set up an MCP server](./cli/tutorials/mcp-setup.md):** Set up an MCP
+ server.
+- **[Automate tasks](./cli/tutorials/automation.md):** Automate tasks.
+
+## Features
+
+Technical documentation for each capability of Gemini CLI.
+
+- **[Extensions](./extensions/index.md):** Extend Gemini CLI with new tools and
+ capabilities.
+- **[Agent Skills](./cli/skills.md):** Use specialized agents for specific
+ tasks.
+- **[Checkpointing](./cli/checkpointing.md):** Automatic session snapshots.
+- **[Headless mode](./cli/headless.md):** Programmatic and scripting interface.
+- **[Hooks](./hooks/index.md):** Customize Gemini CLI behavior with scripts.
+- **[IDE integration](./ide-integration/index.md):** Integrate Gemini CLI with
+ your favorite IDE.
+- **[MCP servers](./tools/mcp-server.md):** Connect to and use remote agents.
+- **[Model routing](./cli/model-routing.md):** Automatic fallback resilience.
+- **[Model selection](./cli/model.md):** Choose the best model for your needs.
+- **[Plan mode π¬](./cli/plan-mode.md):** Use a safe, read-only mode for
+ planning complex changes.
+- **[Subagents π¬](./core/subagents.md):** Using specialized agents for specific
+ tasks.
+- **[Remote subagents π¬](./core/remote-agents.md):** Connecting to and using
+ remote agents.
+- **[Rewind](./cli/rewind.md):** Rewind and replay sessions.
+- **[Sandboxing](./cli/sandbox.md):** Isolate tool execution.
+- **[Settings](./cli/settings.md):** Full configuration reference.
+- **[Telemetry](./cli/telemetry.md):** Usage and performance metric details.
+- **[Token caching](./cli/token-caching.md):** Performance optimization.
+
+## Configuration
+
+Settings and customization options for Gemini CLI.
+
+- **[Custom commands](./cli/custom-commands.md):** Personalized shortcuts.
+- **[Enterprise configuration](./cli/enterprise.md):** Professional environment
+ controls.
+- **[Ignore files (.geminiignore)](./cli/gemini-ignore.md):** Exclusion pattern
+ reference.
+- **[Model configuration](./cli/generation-settings.md):** Fine-tune generation
+ parameters like temperature and thinking budget.
+- **[Project context (GEMINI.md)](./cli/gemini-md.md):** Technical hierarchy of
+ context files.
+- **[System prompt override](./cli/system-prompt.md):** Instruction replacement
+ logic.
+- **[Themes](./cli/themes.md):** UI personalization technical guide.
+- **[Trusted folders](./cli/trusted-folders.md):** Security permission logic.
+
+## Reference
+
+Deep technical documentation and API specifications.
+
+- **[Command reference](./reference/commands.md):** Detailed slash command
+ guide.
+- **[Configuration reference](./reference/configuration.md):** Settings and
+ environment variables.
+- **[Keyboard shortcuts](./reference/keyboard-shortcuts.md):** Productivity
+ tips.
+- **[Memory import processor](./reference/memport.md):** How Gemini CLI
+ processes memory from various sources.
+- **[Policy engine](./reference/policy-engine.md):** Fine-grained execution
+ control.
+- **[Tools reference](./reference/tools.md):** Information on how tools are
+ defined, registered, and used.
+
+## Resources
+
+Support, release history, and legal information.
+
+- **[FAQ](./resources/faq.md):** Answers to frequently asked questions.
+- **[Quota and pricing](./resources/quota-and-pricing.md):** Limits and billing
+ details.
+- **[Terms and privacy](./resources/tos-privacy.md):** Official notices and
+ terms.
+- **[Troubleshooting](./resources/troubleshooting.md):** Common issues and
+ solutions.
+- **[Uninstall](./resources/uninstall.md):** How to uninstall Gemini CLI.
+
+## Development
+
+- **[Contribution guide](/docs/contributing):** How to contribute to Gemini CLI.
+- **[Integration testing](./integration-tests.md):** Running integration tests.
+- **[Issue and PR automation](./issue-and-pr-automation.md):** Automation for
+ issues and pull requests.
+- **[Local development](./local-development.md):** Setting up a local
+ development environment.
+- **[NPM package structure](./npm.md):** The structure of the NPM packages.
+
+## Releases
+
+- **[Release notes](./changelogs/index.md):** Release notes for all versions.
+- **[Stable release](./changelogs/latest.md):** The latest stable release.
+- **[Preview release](./changelogs/preview.md):** The latest preview release.
diff --git a/docs/integration-tests.md b/docs/integration-tests.md
new file mode 100644
index 0000000000000000000000000000000000000000..06ac3a347f2600d8581711fe1de906d038e7c7e1
--- /dev/null
+++ b/docs/integration-tests.md
@@ -0,0 +1,293 @@
+# Integration tests
+
+This document provides information about the integration testing framework used
+in this project.
+
+## Overview
+
+The integration tests are designed to validate the end-to-end functionality of
+Gemini CLI. They execute the built binary in a controlled environment and verify
+that it behaves as expected when interacting with the file system.
+
+These tests are located in the `integration-tests` directory and are run using a
+custom test runner.
+
+## Building the tests
+
+Prior to running any integration tests, you need to create a release bundle that
+you want to actually test:
+
+```bash
+npm run bundle
+```
+
+You must re-run this command after making any changes to the CLI source code,
+but not after making changes to tests.
+
+## Running the tests
+
+The integration tests are not run as part of the default `npm run test` command.
+They must be run explicitly using the `npm run test:integration:all` script.
+
+The integration tests can also be run using the following shortcut:
+
+```bash
+npm run test:e2e
+```
+
+## Running a specific set of tests
+
+To run a subset of test files, you can use
+`npm run ....` where <integration
+test command> is either `test:e2e` or `test:integration*` and ``
+is any of the `.test.js` files in the `integration-tests/` directory. For
+example, the following command runs `list_directory.test.js` and
+`write_file.test.js`:
+
+```bash
+npm run test:e2e list_directory write_file
+```
+
+### Running a single test by name
+
+To run a single test by its name, use the `--test-name-pattern` flag:
+
+```bash
+npm run test:e2e -- --test-name-pattern "reads a file"
+```
+
+### Regenerating model responses
+
+Some integration tests use faked out model responses, which may need to be
+regenerated from time to time as the implementations change.
+
+To regenerate these golden files, set the REGENERATE_MODEL_GOLDENS environment
+variable to "true" when running the tests, for example:
+
+**WARNING**: If running locally you should review these updated responses for
+any information about yourself or your system that gemini may have included in
+these responses.
+
+```bash
+REGENERATE_MODEL_GOLDENS="true" npm run test:e2e
+```
+
+**WARNING**: Make sure you run **await rig.cleanup()** at the end of your test,
+else the golden files will not be updated.
+
+### Deflaking a test
+
+Before adding a **new** integration test, you should test it at least 5 times
+with the deflake script or workflow to make sure that it is not flaky.
+
+### Deflake script
+
+```bash
+npm run deflake -- --runs=5 --command="npm run test:e2e -- -- --test-name-pattern ''"
+```
+
+#### Deflake workflow
+
+```bash
+gh workflow run deflake.yml --ref -f test_name_pattern=""
+```
+
+### Running all tests
+
+To run the entire suite of integration tests, use the following command:
+
+```bash
+npm run test:integration:all
+```
+
+### Sandbox matrix
+
+The `all` command will run tests for `no sandboxing`, `docker` and `podman`.
+Each individual type can be run using the following commands:
+
+```bash
+npm run test:integration:sandbox:none
+```
+
+```bash
+npm run test:integration:sandbox:docker
+```
+
+```bash
+npm run test:integration:sandbox:podman
+```
+
+## Memory regression tests
+
+Memory regression tests are designed to detect heap growth and leaks across key
+CLI scenarios. They are located in the `memory-tests` directory.
+
+These tests are distinct from standard integration tests because they measure
+memory usage and compare it against committed baselines.
+
+### Running memory tests
+
+Memory tests are not run as part of the default `npm run test` or
+`npm run test:e2e` commands. They are run nightly in CI but can be run manually:
+
+```bash
+npm run test:memory
+```
+
+### Updating baselines
+
+If you intentionally change behavior that affects memory usage, you may need to
+update the baselines. Set the `UPDATE_MEMORY_BASELINES` environment variable to
+`true`:
+
+```bash
+UPDATE_MEMORY_BASELINES=true npm run test:memory
+```
+
+This will run the tests, take median snapshots, and overwrite
+`memory-tests/baselines.json`. You should review the changes and commit the
+updated baseline file.
+
+### How it works
+
+The harness (`MemoryTestHarness` in `packages/test-utils`):
+
+- Forces garbage collection multiple times to reduce noise.
+- Takes median snapshots to filter spikes.
+- Compares against baselines with a 10% tolerance.
+- Can analyze sustained leaks across 3 snapshots using `analyzeSnapshots()`.
+
+## Performance regression tests
+
+Performance regression tests are designed to detect wall-clock time, CPU usage,
+and event loop delay regressions across key CLI scenarios. They are located in
+the `perf-tests` directory.
+
+These tests are distinct from standard integration tests because they measure
+performance metrics and compare it against committed baselines.
+
+### Running performance tests
+
+Performance tests are not run as part of the default `npm run test` or
+`npm run test:e2e` commands. They are run nightly in CI but can be run manually:
+
+```bash
+npm run test:perf
+```
+
+### Updating baselines
+
+If you intentionally change behavior that affects performance, you may need to
+update the baselines. Set the `UPDATE_PERF_BASELINES` environment variable to
+`true`:
+
+```bash
+UPDATE_PERF_BASELINES=true npm run test:perf
+```
+
+This will run the tests multiple times (with warmup), apply IQR outlier
+filtering, and overwrite `perf-tests/baselines.json`. You should review the
+changes and commit the updated baseline file.
+
+### How it works
+
+The harness (`PerfTestHarness` in `packages/test-utils`):
+
+- Measures wall-clock time using `performance.now()`.
+- Measures CPU usage using `process.cpuUsage()`.
+- Monitors event loop delay using `perf_hooks.monitorEventLoopDelay()`.
+- Applies IQR (Interquartile Range) filtering to remove outlier samples.
+- Compares against baselines with a 15% tolerance.
+
+## Diagnostics
+
+The integration test runner provides several options for diagnostics to help
+track down test failures.
+
+### Keeping test output
+
+You can preserve the temporary files created during a test run for inspection.
+This is useful for debugging issues with file system operations.
+
+To keep the test output set the `KEEP_OUTPUT` environment variable to `true`.
+
+```bash
+KEEP_OUTPUT=true npm run test:integration:sandbox:none
+```
+
+When output is kept, the test runner will print the path to the unique directory
+for the test run.
+
+### Verbose output
+
+For more detailed debugging, set the `VERBOSE` environment variable to `true`.
+
+```bash
+VERBOSE=true npm run test:integration:sandbox:none
+```
+
+When using `VERBOSE=true` and `KEEP_OUTPUT=true` in the same command, the output
+is streamed to the console and also saved to a log file within the test's
+temporary directory.
+
+The verbose output is formatted to clearly identify the source of the logs:
+
+```
+--- TEST: : ---
+... output from the gemini command ...
+--- END TEST: : ---
+```
+
+## Linting and formatting
+
+To ensure code quality and consistency, the integration test files are linted as
+part of the main build process. You can also manually run the linter and
+auto-fixer.
+
+### Running the linter
+
+To check for linting errors, run the following command:
+
+```bash
+npm run lint
+```
+
+You can include the `:fix` flag in the command to automatically fix any fixable
+linting errors:
+
+```bash
+npm run lint:fix
+```
+
+## Directory structure
+
+The integration tests create a unique directory for each test run inside the
+`.integration-tests` directory. Within this directory, a subdirectory is created
+for each test file, and within that, a subdirectory is created for each
+individual test case.
+
+This structure makes it easy to locate the artifacts for a specific test run,
+file, or case.
+
+```
+.integration-tests/
+βββ /
+ βββ .test.js/
+ βββ /
+ βββ output.log
+ βββ ...other test artifacts...
+```
+
+## Continuous integration
+
+To ensure the integration tests are always run, a GitHub Actions workflow is
+defined in `.github/workflows/chained_e2e.yml`. This workflow automatically runs
+the integrations tests for pull requests against the `main` branch, or when a
+pull request is added to a merge queue.
+
+The workflow runs the tests in different sandboxing environments to ensure
+Gemini CLI is tested across each:
+
+- `sandbox:none`: Runs the tests without any sandboxing.
+- `sandbox:docker`: Runs the tests in a Docker container.
+- `sandbox:podman`: Runs the tests in a Podman container.
diff --git a/docs/issue-and-pr-automation.md b/docs/issue-and-pr-automation.md
new file mode 100644
index 0000000000000000000000000000000000000000..cdcfbb75c78ccf903dba8bdd1b5aff4198b03594
--- /dev/null
+++ b/docs/issue-and-pr-automation.md
@@ -0,0 +1,203 @@
+# Automation and triage processes
+
+This document provides a detailed overview of the automated processes we use to
+manage and triage issues and pull requests. Our goal is to provide prompt
+feedback and ensure that contributions are reviewed and integrated efficiently.
+Understanding this automation will help you as a contributor know what to expect
+and how to best interact with our repository bots.
+
+## Guiding principle: Issues and pull requests
+
+First and foremost, almost every Pull Request (PR) should be linked to a
+corresponding Issue. The issue describes the "what" and the "why" (the bug or
+feature), while the PR is the "how" (the implementation). This separation helps
+us track work, prioritize features, and maintain clear historical context. Our
+automation is built around this principle.
+
+
+> [!NOTE]
+> Issues tagged as "πMaintainers only" are reserved for project
+> maintainers. We will not accept pull requests related to these issues.
+
+---
+
+## Detailed automation workflows
+
+Here is a breakdown of the specific automation workflows that run in our
+repository.
+
+### 1. When you open an issue: `Automated Issue Triage`
+
+This is the first bot you will interact with when you create an issue. Its job
+is to perform an initial analysis and apply the correct labels.
+
+- **Workflow File**: `.github/workflows/gemini-automated-issue-triage.yml`
+- **When it runs**: Immediately after an issue is created or reopened.
+- **What it does**:
+ - It uses a Gemini model to analyze the issue's title and body against a
+ detailed set of guidelines.
+ - **Applies one `area/*` label**: Categorizes the issue into a functional area
+ of the project (for example, `area/ux`, `area/models`, `area/platform`).
+ - **Applies one `kind/*` label**: Identifies the type of issue (for example,
+ `kind/bug`, `kind/enhancement`, `kind/question`).
+ - **Applies one `priority/*` label**: Assigns a priority from P0 (critical) to
+ P3 (low) based on the described impact.
+ - **May apply `status/need-information`**: If the issue lacks critical details
+ (like logs or reproduction steps), it will be flagged for more information.
+ - **May apply `status/need-retesting`**: If the issue references a CLI version
+ that is more than six versions old, it will be flagged for retesting on a
+ current version.
+- **What you should do**:
+ - Fill out the issue template as completely as possible. The more detail you
+ provide, the more accurate the triage will be.
+ - If the `status/need-information` label is added, provide the requested
+ details in a comment.
+
+### 2. When you open a pull request: `Continuous Integration (CI)`
+
+This workflow ensures that all changes meet our quality standards before they
+can be merged.
+
+- **Workflow File**: `.github/workflows/ci.yml`
+- **When it runs**: On every push to a pull request.
+- **What it does**:
+ - **Lint**: Checks that your code adheres to our project's formatting and
+ style rules.
+ - **Test**: Runs our full suite of automated tests across macOS, Windows, and
+ Linux, and on multiple Node.js versions. This is the most time-consuming
+ part of the CI process.
+ - **Post Coverage Comment**: After all tests have successfully passed, a bot
+ will post a comment on your PR. This comment provides a summary of how well
+ your changes are covered by tests.
+- **What you should do**:
+ - Ensure all CI checks pass. A green checkmark β will appear next to your
+ commit when everything is successful.
+ - If a check fails (a red "X" β), click the "Details" link next to the failed
+ check to view the logs, identify the problem, and push a fix.
+
+### 3. Ongoing triage for pull requests: `PR Auditing and Label Sync`
+
+This workflow runs periodically to ensure all open PRs are correctly linked to
+issues and have consistent labels.
+
+- **Workflow File**: `.github/workflows/gemini-scheduled-pr-triage.yml`
+- **When it runs**: Every 15 minutes on all open pull requests.
+- **What it does**:
+ - **Checks for a linked issue**: The bot scans your PR description for a
+ keyword that links it to an issue (for example, `Fixes #123`,
+ `Closes #456`).
+ - **Adds `status/need-issue`**: If no linked issue is found, the bot will add
+ the `status/need-issue` label to your PR. This is a clear signal that an
+ issue needs to be created and linked.
+ - **Synchronizes labels**: If an issue _is_ linked, the bot ensures the PR's
+ labels perfectly match the issue's labels. It will add any missing labels
+ and remove any that don't belong, and it will remove the `status/need-issue`
+ label if it was present.
+- **What you should do**:
+ - **Always link your PR to an issue.** This is the most important step. Add a
+ line like `Resolves #` to your PR description.
+ - This will ensure your PR is correctly categorized and moves through the
+ review process smoothly.
+
+### 4. Ongoing triage for issues: `Scheduled Issue Triage`
+
+This is a fallback workflow to ensure that no issue gets missed by the triage
+process.
+
+- **Workflow File**: `.github/workflows/gemini-scheduled-issue-triage.yml`
+- **When it runs**: Every hour on all open issues.
+- **What it does**:
+ - It actively seeks out issues that either have no labels at all or still have
+ the `status/need-triage` label.
+ - It then triggers the same powerful Gemini-based analysis as the initial
+ triage bot to apply the correct labels.
+- **What you should do**:
+ - You typically don't need to do anything. This workflow is a safety net to
+ ensure every issue is eventually categorized, even if the initial triage
+ fails.
+
+### 5. Automatic unassignment of inactive contributors: `Unassign Inactive Issue Assignees`
+
+To keep the list of open `help wanted` issues accessible to all contributors,
+this workflow automatically removes **external contributors** who have not
+opened a linked pull request within **7 days** of being assigned. Maintainers,
+org members, and repo collaborators with write access or above are always exempt
+and will never be auto-unassigned.
+
+- **Workflow File**: `.github/workflows/unassign-inactive-assignees.yml`
+- **When it runs**: Every day at 09:00 UTC, and can be triggered manually with
+ an optional `dry_run` mode.
+- **What it does**:
+ 1. Finds every open issue labeled `help wanted` that has at least one
+ assignee.
+ 2. Identifies privileged users (team members, repo collaborators with write+
+ access, maintainers) and skips them entirely.
+ 3. For each remaining (external) assignee it reads the issue's timeline to
+ determine:
+ - The exact date they were assigned (using `assigned` timeline events).
+ - Whether they have opened a PR that is already linked/cross-referenced to
+ the issue.
+ 4. Each cross-referenced PR is fetched to verify it is **ready for review**:
+ open and non-draft, or already merged. Draft PRs do not count.
+ 5. If an assignee has been assigned for **more than 7 days** and no qualifying
+ PR is found, they are automatically unassigned and a comment is posted
+ explaining the reason and how to re-claim the issue.
+ 6. Assignees who have a non-draft, open or merged PR linked to the issue are
+ **never** unassigned by this workflow.
+- **What you should do**:
+ - **Open a real PR, not a draft**: Within 7 days of being assigned, open a PR
+ that is ready for review and include `Fixes #` in the
+ description. Draft PRs do not satisfy the requirement and will not prevent
+ auto-unassignment.
+ - **Re-assign if unassigned by mistake**: Comment `/assign` on the issue to
+ assign yourself again.
+ - **Unassign yourself** if you can no longer work on the issue by commenting
+ `/unassign`, so other contributors can pick it up right away.
+
+### 6. Automatically label PRs by size: `PR Size Labeler`
+
+To help maintainers estimate review effort and keep the PR history clean, this
+workflow automatically tags every pull request with a size label representing
+the total volume of line changes.
+
+- **Workflow File**: `.github/workflows/pr-size-labeler.yml`
+- **When it runs**: Immediately after a pull request is created, synchronized
+ (new commits pushed), or reopened. It can also be triggered manually via
+ `workflow_dispatch` with a PR number.
+- **What it does**:
+ - **Calculates total changes**: Summarizes additions and deletions across all
+ changed files in a single consolidated API request.
+ - **Applies standard size labels**:
+ - `size/XS`: < 10 lines changed
+ - `size/S`: 10-49 lines changed
+ - `size/M`: 50-249 lines changed
+ - `size/L`: 250-999 lines changed
+ - `size/XL`: >= 1000 lines changed
+ - **Updates size tag atomically**: Adds the new correct size label and removes
+ any obsolete size labels in one atomic step.
+ - **Updates/Posts PR size info comment**: Instead of spamming a new comment on
+ every commit push, it updates the existing size labeler status comment
+ inline to keep the PR conversation timeline perfectly neat and clean.
+- **What you should do**:
+ - You do not need to take any actions. The workflow runs automatically and
+ updates the label and comment seamlessly as you push new updates.
+
+### 7. Release automation
+
+This workflow handles the process of packaging and publishing new versions of
+Gemini CLI.
+
+- **Workflow File**: `.github/workflows/release-manual.yml`
+- **When it runs**: On a daily schedule for "nightly" releases, and manually for
+ official patch/minor releases.
+- **What it does**:
+ - Automatically builds the project, bumps the version numbers, and publishes
+ the packages to npm.
+ - Creates a corresponding release on GitHub with generated release notes.
+- **What you should do**:
+ - As a contributor, you don't need to do anything for this process. You can be
+ confident that once your PR is merged into the `main` branch, your changes
+ will be included in the very next nightly release.
+
+We hope this detailed overview is helpful. If you have any questions about our
+automation or processes, don't hesitate to ask!
diff --git a/docs/local-development.md b/docs/local-development.md
new file mode 100644
index 0000000000000000000000000000000000000000..e6f862044dc01da6b0b6b86007d5400c11208658
--- /dev/null
+++ b/docs/local-development.md
@@ -0,0 +1,182 @@
+# Local development guide
+
+This guide provides instructions for setting up and using local development
+features for Gemini CLI.
+
+## Tracing
+
+Gemini CLI uses OpenTelemetry (OTel) to record traces that help you debug agent
+behavior. Traces instrument key events like model calls, tool scheduler
+operations, and tool calls.
+
+Traces provide deep visibility into agent behavior and help you debug complex
+issues. They are captured automatically when you enable telemetry.
+
+### View traces
+
+You can view traces using Genkit Developer UI, Jaeger, or Google Cloud.
+
+#### Use Genkit
+
+Genkit provides a web-based UI for viewing traces and other telemetry data.
+
+1. **Start the Genkit telemetry server:**
+
+ Run the following command to start the Genkit server:
+
+ ```bash
+ npm run telemetry -- --target=genkit
+ ```
+
+ The script will output the URL for the Genkit Developer UI. For example:
+ `Genkit Developer UI: http://localhost:4000`
+
+2. **Run Gemini CLI:**
+
+ In a separate terminal, run your Gemini CLI command:
+
+ ```bash
+ gemini
+ ```
+
+3. **View the traces:**
+
+ Open the Genkit Developer UI URL in your browser and navigate to the
+ **Traces** tab to view the traces.
+
+#### Use Jaeger
+
+You can view traces in the Jaeger UI for local development.
+
+1. **Start the telemetry collector:**
+
+ Run the following command in your terminal to download and start Jaeger and
+ an OTel collector:
+
+ ```bash
+ npm run telemetry -- --target=local
+ ```
+
+ This command configures your workspace for local telemetry and provides a
+ link to the Jaeger UI (usually `http://localhost:16686`).
+
+ - **Collector logs:** `~/.gemini/tmp//otel/collector.log`
+
+2. **Run Gemini CLI:**
+
+ In a separate terminal, run your Gemini CLI command:
+
+ ```bash
+ gemini
+ ```
+
+3. **View the traces:**
+
+ After running your command, open the Jaeger UI link in your browser to view
+ the traces.
+
+#### Use Google Cloud
+
+You can use an OpenTelemetry collector to forward telemetry data to Google Cloud
+Trace for custom processing or routing.
+
+
+> [!WARNING]
+> Ensure you complete the
+> [Google Cloud telemetry prerequisites](./cli/telemetry.md#prerequisites)
+> (Project ID, authentication, IAM roles, and APIs) before using this method.
+
+1. **Configure `.gemini/settings.json`:**
+
+ ```json
+ {
+ "telemetry": {
+ "enabled": true,
+ "target": "gcp",
+ "useCollector": true
+ }
+ }
+ ```
+
+2. **Start the telemetry collector:**
+
+ Run the following command to start a local OTel collector that forwards to
+ Google Cloud:
+
+ ```bash
+ npm run telemetry -- --target=gcp
+ ```
+
+ The script outputs links to view traces, metrics, and logs in the Google
+ Cloud Console.
+
+ - **Collector logs:** `~/.gemini/tmp//otel/collector-gcp.log`
+
+3. **Run Gemini CLI:**
+
+ In a separate terminal, run your Gemini CLI command:
+
+ ```bash
+ gemini
+ ```
+
+4. **View logs, metrics, and traces:**
+
+ After sending prompts, view your data in the Google Cloud Console. See the
+ [telemetry documentation](./cli/telemetry.md#view-google-cloud-telemetry)
+ for links to Logs, Metrics, and Trace explorers.
+
+For more detailed information on telemetry, see the
+[telemetry documentation](./cli/telemetry.md).
+
+### Instrument code with traces
+
+You can add traces to your own code for more detailed instrumentation.
+
+Adding traces helps you debug and understand the flow of execution. Use the
+`runInDevTraceSpan` function to wrap any section of code in a trace span.
+
+Here is a basic example:
+
+```typescript
+import { runInDevTraceSpan } from '@google/gemini-cli-core';
+import { GeminiCliOperation } from '@google/gemini-cli-core/lib/telemetry/constants.js';
+
+await runInDevTraceSpan(
+ {
+ operation: GeminiCliOperation.ToolCall,
+ attributes: {
+ [GEN_AI_AGENT_NAME]: 'gemini-cli',
+ },
+ },
+ async ({ metadata }) => {
+ // metadata allows you to record the input and output of the
+ // operation as well as other attributes.
+ metadata.input = { key: 'value' };
+ // Set custom attributes.
+ metadata.attributes['custom.attribute'] = 'custom.value';
+
+ // Your code to be traced goes here.
+ try {
+ const output = await somethingRisky();
+ metadata.output = output;
+ return output;
+ } catch (e) {
+ metadata.error = e;
+ throw e;
+ }
+ },
+);
+```
+
+In this example:
+
+- `operation`: The operation type of the span, represented by the
+ `GeminiCliOperation` enum.
+- `metadata.input`: (Optional) An object containing the input data for the
+ traced operation.
+- `metadata.output`: (Optional) An object containing the output data from the
+ traced operation.
+- `metadata.attributes`: (Optional) A record of custom attributes to add to the
+ span.
+- `metadata.error`: (Optional) An error object to record if the operation fails.
diff --git a/docs/npm.md b/docs/npm.md
new file mode 100644
index 0000000000000000000000000000000000000000..3ceab3c5e717d9c15123eccfd1682839339c53d1
--- /dev/null
+++ b/docs/npm.md
@@ -0,0 +1,62 @@
+# Package overview
+
+This monorepo contains two main packages: `@google/gemini-cli` and
+`@google/gemini-cli-core`.
+
+## `@google/gemini-cli`
+
+This is the main package for Gemini CLI. It is responsible for the user
+interface, command parsing, and all other user-facing functionality.
+
+When this package is published, it is bundled into a single executable file.
+This bundle includes all of the package's dependencies, including
+`@google/gemini-cli-core`. This means that whether a user installs the package
+with `npm install -g @google/gemini-cli` or runs it directly with
+`npx @google/gemini-cli`, they are using this single, self-contained executable.
+
+## `@google/gemini-cli-core`
+
+This package contains the core logic for interacting with the Gemini API. It is
+responsible for making API requests, handling authentication, and managing the
+local cache.
+
+This package is not bundled. When it is published, it is published as a standard
+Node.js package with its own dependencies. This allows it to be used as a
+standalone package in other projects, if needed. All transpiled js code in the
+`dist` folder is included in the package.
+
+## NPM workspaces
+
+This project uses
+[NPM Workspaces](https://docs.npmjs.com/cli/v10/using-npm/workspaces) to manage
+the packages within this monorepo. This simplifies development by allowing us to
+manage dependencies and run scripts across multiple packages from the root of
+the project.
+
+### How it works
+
+The root `package.json` file defines the workspaces for this project:
+
+```json
+{
+ "workspaces": ["packages/*"]
+}
+```
+
+This tells NPM that any folder inside the `packages` directory is a separate
+package that should be managed as part of the workspace.
+
+### Benefits of workspaces
+
+- **Simplified dependency management**: Running `npm install` from the root of
+ the project will install all dependencies for all packages in the workspace
+ and link them together. This means you don't need to run `npm install` in each
+ package's directory.
+- **Automatic linking**: Packages within the workspace can depend on each other.
+ When you run `npm install`, NPM will automatically create symlinks between the
+ packages. This means that when you make changes to one package, the changes
+ are immediately available to other packages that depend on it.
+- **Simplified script execution**: You can run scripts in any package from the
+ root of the project using the `--workspace` flag. For example, to run the
+ `build` script in the `cli` package, you can run
+ `npm run build --workspace @google/gemini-cli`.
diff --git a/docs/redirects.json b/docs/redirects.json
new file mode 100644
index 0000000000000000000000000000000000000000..db2dae4333c22501a2f29b3fd4531ba585db9b83
--- /dev/null
+++ b/docs/redirects.json
@@ -0,0 +1,21 @@
+{
+ "/docs/architecture": "/docs/cli/index",
+ "/docs/cli/commands": "/docs/reference/commands",
+ "/docs/cli": "/docs",
+ "/docs/cli/index": "/docs",
+ "/docs/cli/keyboard-shortcuts": "/docs/reference/keyboard-shortcuts",
+ "/docs/cli/uninstall": "/docs/resources/uninstall",
+ "/docs/core/concepts": "/docs",
+ "/docs/core/memport": "/docs/reference/memport",
+ "/docs/core/policy-engine": "/docs/reference/policy-engine",
+ "/docs/core/tools-api": "/docs/reference/tools",
+ "/docs/reference/tools-api": "/docs/reference/tools",
+ "/docs/faq": "/docs/resources/faq",
+ "/docs/get-started/configuration": "/docs/reference/configuration",
+ "/docs/get-started/configuration-v1": "/docs/reference/configuration",
+ "/docs/get-started/examples": "/docs/get-started/index",
+ "/docs/index": "/docs",
+ "/docs/quota-and-pricing": "/docs/resources/quota-and-pricing",
+ "/docs/tos-privacy": "/docs/resources/tos-privacy",
+ "/docs/troubleshooting": "/docs/resources/troubleshooting"
+}
diff --git a/docs/release-confidence.md b/docs/release-confidence.md
new file mode 100644
index 0000000000000000000000000000000000000000..7b6bd06249049075c3ddc7f0a7026f9c77ec8431
--- /dev/null
+++ b/docs/release-confidence.md
@@ -0,0 +1,168 @@
+# Release confidence strategy
+
+This document outlines the strategy for gaining confidence in every release of
+Gemini CLI. It serves as a checklist and quality gate for release manager to
+ensure we are shipping a high-quality product.
+
+## The goal
+
+To answer the question, "Is this release _truly_ ready for our users?" with a
+high degree of confidence, based on a holistic evaluation of automated signals,
+manual verification, and data.
+
+## Level 1: Automated gates (must pass)
+
+These are the baseline requirements. If any of these fail, the release is a
+no-go.
+
+### 1. CI/CD health
+
+All workflows in `.github/workflows/ci.yml` must pass on the `main` branch (for
+nightly) or the release branch (for preview/stable).
+
+- **Platforms:** Tests must pass on **Linux and macOS**.
+
+- **Checks:**
+ - **Linting:** No linting errors (ESLint, Prettier, etc.).
+ - **Typechecking:** No TypeScript errors.
+ - **Unit Tests:** All unit tests in `packages/core` and `packages/cli` must
+ pass.
+ - **Build:** The project must build and bundle successfully.
+
+### 2. End-to-end (E2E) tests
+
+All workflows in `.github/workflows/chained_e2e.yml` must pass.
+
+- **Platforms:** **Linux, macOS and Windows**.
+- **Sandboxing:** Tests must pass with both `sandbox:none` and `sandbox:docker`
+ on Linux.
+
+### 3. Post-deployment smoke tests
+
+After a release is published to npm, the `smoke-test.yml` workflow runs. This
+must pass to confirm the package is installable and the binary is executable.
+
+- **Command:** `npx -y @google/gemini-cli@ --version` must return the
+ correct version without error.
+- **Platform:** Currently runs on `ubuntu-latest`.
+
+## Level 2: Manual verification and dogfooding
+
+Automated tests cannot catch everything, especially UX issues.
+
+### 1. Dogfooding via `preview` tag
+
+The weekly release cadence promotes code from `main` -> `nightly` -> `preview`
+-> `stable`.
+
+- **Requirement:** The `preview` release must be used by maintainers for at
+ least **one week** before being promoted to `stable`.
+- **Action:** Maintainers should install the preview version locally:
+ ```bash
+ npm install -g @google/gemini-cli@preview
+ ```
+- **Goal:** To catch regressions and UX issues in day-to-day usage before they
+ reach the broad user base.
+
+### 2. Critical user journey (CUJ) checklist
+
+Before promoting a `preview` release to `stable`, a release manager must
+manually run through this checklist.
+
+- **Setup:**
+
+ - [ ] Uninstall any existing global version:
+ `npm uninstall -g @google/gemini-cli`
+ - [ ] Clear npx cache (optional but recommended): `npm cache clean --force`
+ - [ ] Install the preview version: `npm install -g @google/gemini-cli@preview`
+ - [ ] Verify version: `gemini --version`
+
+- **Authentication:**
+
+ - [ ] In interactive mode run `/auth` and verify all sign in flows work:
+ - [ ] Sign in with Google
+ - [ ] API Key
+ - [ ] Vertex AI
+
+- **Basic prompting:**
+
+ - [ ] Run `gemini "Tell me a joke"` and verify a sensible response.
+ - [ ] Run in interactive mode: `gemini`. Ask a follow-up question to test
+ context.
+
+- **Piped input:**
+
+ - [ ] Run `echo "Summarize this" | gemini` and verify it processes stdin.
+
+- **Context management:**
+
+ - [ ] In interactive mode, use `@file` to add a local file to context. Ask a
+ question about it.
+
+- **Settings:**
+
+ - [ ] In interactive mode run `/settings` and make modifications
+ - [ ] Validate that setting is changed
+
+- **Function calling:**
+ - [ ] In interactive mode, ask gemini to "create a file named hello.md with
+ the content 'hello world'" and verify the file is created correctly.
+
+If any of these CUJs fail, the release is a no-go until a patch is applied to
+the `preview` channel.
+
+### 3. Pre-Launch bug bash (tier 1 and 2 launches)
+
+For high-impact releases, an organized bug bash is required to ensure a higher
+level of quality and to catch issues across a wider range of environments and
+use cases.
+
+**Definition of tiers:**
+
+- **Tier 1:** Industry-Moving News π
+- **Tier 2:** Important News for Our Users π£
+- **Tier 3:** Relevant, but Not Life-Changing π‘
+- **Tier 4:** Bug Fixes βοΈ
+
+**Requirement:**
+
+A bug bash must be scheduled at least **72 hours in advance** of any Tier 1 or
+Tier 2 launch.
+
+**Rule of thumb:**
+
+A bug bash should be considered for any release that involves:
+
+- A blog post
+- Coordinated social media announcements
+- Media relations or press outreach
+- A "Turbo" launch event
+
+## Level 3: Telemetry and data review
+
+### Dashboard health
+
+- [ ] Go to `go/gemini-cli-dash`.
+- [ ] Navigate to the "Tool Call" tab.
+- [ ] Validate that there are no spikes in errors for the release you would like
+ to promote.
+
+### Model evaluation
+
+- [ ] Navigate to `go/gemini-cli-offline-evals-dash`.
+- [ ] Make sure that the release you want to promote's recurring run is within
+ average eval runs.
+
+## The "go/no-go" decision
+
+Before triggering the `Release: Promote` workflow to move `preview` to `stable`:
+
+1. [ ] **Level 1:** CI and E2E workflows are green for the commit corresponding
+ to the current `preview` tag.
+2. [ ] **Level 2:** The `preview` version has been out for one week, and the
+ CUJ checklist has been completed successfully by a release manager. No
+ blocking issues have been reported.
+3. [ ] **Level 3:** Dashboard Health and Model Evaluation checks have been
+ completed and show no regressions.
+
+If all checks pass, proceed with the promotion.
diff --git a/docs/releases.md b/docs/releases.md
new file mode 100644
index 0000000000000000000000000000000000000000..70a9f069ce2f4fafb331d88f37e9e2f33985037a
--- /dev/null
+++ b/docs/releases.md
@@ -0,0 +1,550 @@
+# Gemini CLI releases
+
+
+> [!IMPORTANT]
+> **Coordinate with the Release Manager:** The release manager is responsible for coordinating patches and releases. Please update them before performing any of the release actions described in this document.
+
+## `dev` vs `prod` environment
+
+Our release flows support both `dev` and `prod` environments.
+
+The `dev` environment pushes to a private GitHub-hosted NPM repository, with the
+package names beginning with `@google-gemini/**` instead of `@google/**`.
+
+The `prod` environment pushes to the public global NPM registry via Wombat
+Dressing Room, which is Google's system for managing NPM packages in the
+`@google/**` namespace. The packages are all named `@google/**`.
+
+More information can be found about these systems in the
+[NPM Package Overview](npm.md)
+
+### Package scopes
+
+| Package | `prod` (Wombat Dressing Room) | `dev` (GitHub Private NPM Repo) |
+| ---------- | ----------------------------- | ----------------------------------------- |
+| CLI | @google/gemini-cli | @google-gemini/gemini-cli |
+| Core | @google/gemini-cli-core | @google-gemini/gemini-cli-core A2A Server |
+| A2A Server | @google/gemini-cli-a2a-server | @google-gemini/gemini-cli-a2a-server |
+
+## Release cadence and tags
+
+We will follow https://semver.org/ as closely as possible but will call out when
+or if we have to deviate from it. Our weekly releases will be minor version
+increments and any bug or hotfixes between releases will go out as patch
+versions on the most recent release.
+
+Each Tuesday ~20:00 UTC new Stable and Preview releases will be cut. The
+promotion flow is:
+
+- Code is committed to main and pushed each night to nightly
+- After no more than 1 week on main, code is promoted to the `preview` channel
+- After 1 week the most recent `preview` channel is promoted to `stable` channel
+- Patch fixes will be produced against both `preview` and `stable` as needed,
+ with the final 'patch' version number incrementing each time.
+
+### Preview
+
+These releases will not have been fully vetted and may contain regressions or
+other outstanding issues. Help us test and install with `preview` tag.
+
+```bash
+npm install -g @google/gemini-cli@preview
+```
+
+### Stable
+
+This will be the full promotion of last week's release + any bug fixes and
+validations. Use `latest` tag.
+
+```bash
+npm install -g @google/gemini-cli@latest
+```
+
+### Nightly
+
+- New releases will be published each day at UTC 00:00. This will be all changes
+ from the main branch as represented at time of release. It should be assumed
+ there are pending validations and issues. Use `nightly` tag.
+
+```bash
+npm install -g @google/gemini-cli@nightly
+```
+
+## Weekly release promotion
+
+Each Tuesday, the on-call engineer will trigger the "Promote Release" workflow.
+This single action automates the entire weekly release process:
+
+1. **Promotes preview to stable:** The workflow identifies the latest `preview`
+ release and promotes it to `stable`. This becomes the new `latest` version
+ on npm.
+2. **Promotes nightly to preview:** The latest `nightly` release is then
+ promoted to become the new `preview` version.
+3. **Prepares for next nightly:** A pull request is automatically created and
+ merged to bump the version in `main` in preparation for the next nightly
+ release.
+
+This process ensures a consistent and reliable release cadence with minimal
+manual intervention.
+
+### Source of truth for versioning
+
+To ensure the highest reliability, the release promotion process uses the **NPM
+registry as the single source of truth** for determining the current version of
+each release channel (`stable`, `preview`, and `nightly`).
+
+1. **Fetch from NPM:** The workflow begins by querying NPM's `dist-tags`
+ (`latest`, `preview`, `nightly`) to get the exact version strings for the
+ packages currently available to users.
+2. **Cross-check for integrity:** For each version retrieved from NPM, the
+ workflow performs a critical integrity check:
+ - It verifies that a corresponding **git tag** exists in the repository.
+ - It verifies that a corresponding **GitHub release** has been created.
+3. **Halt on discrepancy:** If either the git tag or the GitHub Release is
+ missing for a version listed on NPM, the workflow will immediately fail.
+ This strict check prevents promotions from a broken or incomplete previous
+ release and alerts the on-call engineer to a release state inconsistency
+ that must be manually resolved.
+4. **Calculate next version:** Only after these checks pass does the workflow
+ proceed to calculate the next semantic version based on the trusted version
+ numbers retrieved from NPM.
+
+This NPM-first approach, backed by integrity checks, makes the release process
+highly robust and prevents the kinds of versioning discrepancies that can arise
+from relying solely on git history or API outputs.
+
+## Manual releases
+
+For situations requiring a release outside of the regular nightly and weekly
+promotion schedule, and NOT already covered by patching process, you can use the
+`Release: Manual` workflow. This workflow provides a direct way to publish a
+specific version from any branch, tag, or commit SHA.
+
+### How to create a manual release
+
+1. Navigate to the **Actions** tab of the repository.
+2. Select the **Release: Manual** workflow from the list.
+3. Click the **Run workflow** dropdown button.
+4. Fill in the required inputs:
+ - **Version**: The exact version to release (for example, `v0.6.1`). This
+ must be a valid semantic version with a `v` prefix.
+ - **Ref**: The branch, tag, or full commit SHA to release from.
+ - **NPM Channel**: The npm channel to publish to. The options are `preview`,
+ `nightly`, `latest` (for stable releases), and `dev`. The default is
+ `dev`.
+ - **Dry Run**: Leave as `true` to run all steps without publishing, or set
+ to `false` to perform a live release.
+ - **Force Skip Tests**: Set to `true` to skip the test suite. This is not
+ recommended for production releases.
+ - **Skip GitHub Release**: Set to `true` to skip creating a GitHub release
+ and create an npm release only.
+ - **Environment**: Select the appropriate environment. The `dev` environment
+ is intended for testing. The `prod` environment is intended for production
+ releases. `prod` is the default and will require authorization from a
+ release administrator.
+5. Click **Run workflow**.
+
+The workflow will then proceed to test (if not skipped), build, and publish the
+release. If the workflow fails during a non-dry run, it will automatically
+create a GitHub issue with the failure details.
+
+## Rollback/rollforward
+
+In the event that a release has a critical regression, you can quickly roll back
+to a previous stable version or roll forward to a new patch by changing the npm
+`dist-tag`. The `Release: Change Tags` workflow provides a safe and controlled
+way to do this.
+
+This is the preferred method for both rollbacks and rollforwards, as it does not
+require a full release cycle.
+
+### How to change a release tag
+
+1. Navigate to the **Actions** tab of the repository.
+2. Select the **Release: Change Tags** workflow from the list.
+3. Click the **Run workflow** dropdown button.
+4. Fill in the required inputs:
+ - **Version**: The existing package version that you want to point the tag
+ to (for example, `0.5.0-preview-2`). This version **must** already be
+ published to the npm registry.
+ - **Channel**: The npm `dist-tag` to apply (for example, `preview`,
+ `stable`).
+ - **Dry Run**: Leave as `true` to log the action without making changes, or
+ set to `false` to perform the live tag change.
+ - **Environment**: Select the appropriate environment. The `dev` environment
+ is intended for testing. The `prod` environment is intended for production
+ releases. `prod` is the default and will require authorization from a
+ release administrator.
+5. Click **Run workflow**.
+
+The workflow will then run `npm dist-tag add` for the appropriate `gemini-cli`,
+`gemini-cli-core` and `gemini-cli-a2a-server` packages, pointing the specified
+channel to the specified version.
+
+## Patching
+
+If a critical bug that is already fixed on `main` needs to be patched on a
+`stable` or `preview` release, the process is now highly automated.
+
+### How to patch
+
+#### 1. Create the patch pull request
+
+There are two ways to create a patch pull request:
+
+**Option A: From a GitHub comment (recommended)**
+
+After a pull request containing the fix has been merged, a maintainer can add a
+comment on that same PR with the following format:
+
+`/patch [channel]`
+
+- **channel** (optional):
+ - _no channel_ - patches both stable and preview channels (default,
+ recommended for most fixes)
+ - `both` - patches both stable and preview channels (same as default)
+ - `stable` - patches only the stable channel
+ - `preview` - patches only the preview channel
+
+Examples:
+
+- `/patch` (patches both stable and preview - default)
+- `/patch both` (patches both stable and preview - explicit)
+- `/patch stable` (patches only stable)
+- `/patch preview` (patches only preview)
+
+The `Release: Patch from Comment` workflow will automatically find the merge
+commit SHA and trigger the `Release: Patch (1) Create PR` workflow. If the PR is
+not yet merged, it will post a comment indicating the failure.
+
+**Option B: Manually triggering the workflow**
+
+Navigate to the **Actions** tab and run the **Release: Patch (1) Create PR**
+workflow.
+
+- **Commit**: The full SHA of the commit on `main` that you want to cherry-pick.
+- **Channel**: The channel you want to patch (`stable` or `preview`).
+
+This workflow will automatically:
+
+1. Find the latest release tag for the channel.
+2. Create a release branch from that tag if one doesn't exist (for example,
+ `release/v0.5.1-pr-12345`).
+3. Create a new hotfix branch from the release branch.
+4. Cherry-pick your specified commit into the hotfix branch.
+5. Create a pull request from the hotfix branch back to the release branch.
+
+#### 2. Review and merge
+
+Review the automatically created pull request(s) to ensure the cherry-pick was
+successful and the changes are correct. Once approved, merge the pull request.
+
+
+> [!WARNING]
+> The `release/*` branches are protected by branch protection
+> rules. A pull request to one of these branches requires at least one review from
+> a code owner before it can be merged. This ensures that no unauthorized code is
+> released.
+
+#### 2.5. Adding multiple commits to a hotfix (advanced)
+
+If you need to include multiple fixes in a single patch release, you can add
+additional commits to the hotfix branch after the initial patch PR has been
+created:
+
+1. **Start with the primary fix**: Use `/patch` (or `/patch both`) on the most
+ important PR to create the initial hotfix branch and PR.
+
+2. **Checkout the hotfix branch locally**:
+
+ ```bash
+ git fetch origin
+ git checkout hotfix/v0.5.1/stable/cherry-pick-abc1234 # Use the actual branch name from the PR
+ ```
+
+3. **Cherry-pick additional commits**:
+
+ ```bash
+ git cherry-pick
+ git cherry-pick
+ # Add as many commits as needed
+ ```
+
+4. **Push the updated branch**:
+
+ ```bash
+ git push origin hotfix/v0.5.1/stable/cherry-pick-abc1234
+ ```
+
+5. **Test and review**: The existing patch PR will automatically update with
+ your additional commits. Test thoroughly since you're now releasing multiple
+ changes together.
+
+6. **Update the PR description**: Consider updating the PR title and description
+ to reflect that it includes multiple fixes.
+
+This approach lets you group related fixes into a single patch release while
+maintaining full control over what gets included and how conflicts are resolved.
+
+#### 3. Automatic release
+
+Upon merging the pull request, the `Release: Patch (2) Trigger` workflow is
+automatically triggered. It will then start the `Release: Patch (3) Release`
+workflow, which will:
+
+1. Build and test the patched code.
+2. Publish the new patch version to npm.
+3. Create a new GitHub release with the patch notes.
+
+This fully automated process ensures that patches are created and released
+consistently and reliably.
+
+#### Troubleshooting: Older branch workflows
+
+**Issue**: If the patch trigger workflow fails with errors like "Resource not
+accessible by integration" or references to non-existent workflow files (for
+example, `patch-release.yml`), this indicates the hotfix branch contains an
+outdated version of the workflow files.
+
+**Root cause**: When a PR is merged, GitHub Actions runs the workflow definition
+from the **source branch** (the hotfix branch), not from the target branch (the
+release branch). If the hotfix branch was created from an older release branch
+that predates workflow improvements, it will use the old workflow logic.
+
+**Solutions**:
+
+**Option 1: Manual trigger (quick fix)** Manually trigger the updated workflow
+from the branch with the latest workflow code:
+
+```bash
+# For a preview channel patch with tests skipped
+gh workflow run release-patch-2-trigger.yml --ref \
+ --field ref="hotfix/v0.6.0-preview.2/preview/cherry-pick-abc1234" \
+ --field workflow_ref= \
+ --field dry_run=false \
+ --field force_skip_tests=true
+
+# For a stable channel patch
+gh workflow run release-patch-2-trigger.yml --ref \
+ --field ref="hotfix/v0.5.1/stable/cherry-pick-abc1234" \
+ --field workflow_ref= \
+ --field dry_run=false \
+ --field force_skip_tests=false
+
+# Example using main branch (most common case)
+gh workflow run release-patch-2-trigger.yml --ref main \
+ --field ref="hotfix/v0.6.0-preview.2/preview/cherry-pick-abc1234" \
+ --field workflow_ref=main \
+ --field dry_run=false \
+ --field force_skip_tests=true
+```
+
+**Note**: Replace `` with the branch containing
+the latest workflow improvements (usually `main`, but could be a feature branch
+if testing updates).
+
+**Option 2: Update the hotfix branch** Merge the latest main branch into your
+hotfix branch to get the updated workflows:
+
+```bash
+git checkout hotfix/v0.6.0-preview.2/preview/cherry-pick-abc1234
+git merge main
+git push
+```
+
+Then close and reopen the PR to retrigger the workflow with the updated version.
+
+**Option 3: Direct release trigger** Skip the trigger workflow entirely and
+directly run the release workflow:
+
+```bash
+# Replace channel and release_ref with appropriate values
+gh workflow run release-patch-3-release.yml --ref main \
+ --field type="preview" \
+ --field dry_run=false \
+ --field force_skip_tests=true \
+ --field release_ref="release/v0.6.0-preview.2"
+```
+
+### Docker
+
+We also run a Google cloud build called
+[release-docker.yml](../.gcp/release-docker.yml). Which publishes the sandbox
+docker to match your release. This will also be moved to GH and combined with
+the main release file once service account permissions are sorted out.
+
+## Release validation
+
+After pushing a new release smoke testing should be performed to ensure that the
+packages are working as expected. This can be done by installing the packages
+locally and running a set of tests to ensure that they are functioning
+correctly.
+
+- `npx -y @google/gemini-cli@latest --version` to validate the push worked as
+ expected if you were not doing a rc or dev tag
+- `npx -y @google/gemini-cli@ --version` to validate the tag pushed
+ appropriately
+- _This is destructive locally_
+ `npm uninstall @google/gemini-cli && npm uninstall -g @google/gemini-cli && npm cache clean --force && npm install @google/gemini-cli@`
+- Smoke testing a basic run through of exercising a few llm commands and tools
+ is recommended to ensure that the packages are working as expected. We'll
+ codify this more in the future.
+
+## Local testing and validation: Changes to the packaging and publishing process
+
+If you need to test the release process without actually publishing to NPM or
+creating a public GitHub release, you can trigger the workflow manually from the
+GitHub UI.
+
+1. Go to the
+ [Actions tab](https://github.com/google-gemini/gemini-cli/actions/workflows/release-manual.yml)
+ of the repository.
+2. Click on the "Run workflow" dropdown.
+3. Leave the `dry_run` option checked (`true`).
+4. Click the "Run workflow" button.
+
+This will run the entire release process but will skip the `npm publish` and
+`gh release create` steps. You can inspect the workflow logs to ensure
+everything is working as expected.
+
+It is crucial to test any changes to the packaging and publishing process
+locally before committing them. This ensures that the packages will be published
+correctly and that they will work as expected when installed by a user.
+
+To validate your changes, you can perform a dry run of the publishing process.
+This will simulate the publishing process without actually publishing the
+packages to the npm registry.
+
+```bash
+npm_package_version=9.9.9 SANDBOX_IMAGE_REGISTRY="registry" SANDBOX_IMAGE_NAME="thename" npm run publish:npm --dry-run
+```
+
+This command will do the following:
+
+1. Build all the packages.
+2. Run all the prepublish scripts.
+3. Create the package tarballs that would be published to npm.
+4. Print a summary of the packages that would be published.
+
+You can then inspect the generated tarballs to ensure that they contain the
+correct files and that the `package.json` files have been updated correctly. The
+tarballs will be created in the root of each package's directory (for example,
+`packages/cli/google-gemini-cli-0.1.6.tgz`).
+
+By performing a dry run, you can be confident that your changes to the packaging
+process are correct and that the packages will be published successfully.
+
+## Release deep dive
+
+The release process creates two distinct types of artifacts for different
+distribution channels: standard packages for the NPM registry and a single,
+self-contained executable for GitHub Releases.
+
+Here are the key stages:
+
+**Stage 1: Pre-release sanity checks and versioning**
+
+- **What happens:** Before any files are moved, the process ensures the project
+ is in a good state. This involves running tests, linting, and type-checking
+ (`npm run preflight`). The version number in the root `package.json` and
+ `packages/cli/package.json` is updated to the new release version.
+
+**Stage 2: Building the source code for NPM**
+
+- **What happens:** The TypeScript source code in `packages/core/src` and
+ `packages/cli/src` is compiled into standard JavaScript.
+- **File movement:**
+ - `packages/core/src/**/*.ts` -> compiled to -> `packages/core/dist/`
+ - `packages/cli/src/**/*.ts` -> compiled to -> `packages/cli/dist/`
+- **Why:** The TypeScript code written during development needs to be converted
+ into plain JavaScript that can be run by Node.js. The `core` package is built
+ first as the `cli` package depends on it.
+
+**Stage 3: Publishing standard packages to NPM**
+
+- **What happens:** The `npm publish` command is run for the
+ `@google/gemini-cli-core` and `@google/gemini-cli` packages.
+- **Why:** This publishes them as standard Node.js packages. Users installing
+ via `npm install -g @google/gemini-cli` will download these packages, and
+ `npm` will handle installing the `@google/gemini-cli-core` dependency
+ automatically. The code in these packages is not bundled into a single file.
+
+**Stage 4: Assembling and creating the GitHub release asset**
+
+This stage happens _after_ the NPM publish and creates the single-file
+executable that enables `npx` usage directly from the GitHub repository.
+
+1. **The JavaScript bundle is created:**
+
+ - **What happens:** The built JavaScript from both `packages/core/dist` and
+ `packages/cli/dist`, along with all third-party JavaScript dependencies,
+ are bundled by `esbuild` into a single, executable JavaScript file (for
+ example, `gemini.js`). The `node-pty` library is excluded from this bundle
+ as it contains native binaries.
+ - **Why:** This creates a single, optimized file that contains all the
+ necessary application code. It simplifies execution for users who want to
+ run the CLI without a full `npm install`, as all dependencies (including
+ the `core` package) are included directly.
+
+2. **The `bundle` directory is assembled:**
+
+ - **What happens:** A temporary `bundle` folder is created at the project
+ root. The single `gemini.js` executable is placed inside it, along with
+ other essential files.
+ - **File movement:**
+ - `gemini.js` (from esbuild) -> `bundle/gemini.js`
+ - `README.md` -> `bundle/README.md`
+ - `LICENSE` -> `bundle/LICENSE`
+ - `packages/cli/src/utils/*.sb` (sandbox profiles) -> `bundle/`
+ - **Why:** This creates a clean, self-contained directory with everything
+ needed to run the CLI and understand its license and usage.
+
+3. **The GitHub release is created:**
+ - **What happens:** The contents of the `bundle` directory, including the
+ `gemini.js` executable, are attached as assets to a new GitHub Release.
+ - **Why:** This makes the single-file version of the CLI available for
+ direct download and enables the
+ `npx https://github.com/google-gemini/gemini-cli` command, which downloads
+ and runs this specific bundled asset.
+
+**Summary of artifacts**
+
+- **NPM:** Publishes standard, un-bundled Node.js packages. The primary artifact
+ is the code in `packages/cli/dist`, which depends on
+ `@google/gemini-cli-core`.
+- **GitHub release:** Publishes a single, bundled `gemini.js` file that contains
+ all dependencies, for easy execution via `npx`.
+
+This dual-artifact process ensures that both traditional `npm` users and those
+who prefer the convenience of `npx` have an optimized experience.
+
+## Notifications
+
+Failing release workflows will automatically create an issue with the label
+`release-failure`.
+
+A notification will be posted to the maintainer's chat channel when issues with
+this type are created.
+
+### Modifying chat notifications
+
+Notifications use
+[GitHub for Google Chat](https://workspace.google.com/marketplace/app/github_for_google_chat/536184076190).
+To modify the notifications, use `/github-settings` within the chat space.
+
+
+> [!WARNING]
+> The following instructions describe a fragile workaround that depends on the
+> internal structure of the chat application's UI. It is likely to break with
+> future updates.
+
+The list of available labels is not currently populated correctly. If you want
+to add a label that does not appear alphabetically in the first 30 labels in the
+repo, you must use your browser's developer tools to manually modify the UI:
+
+1. Open your browser's developer tools (for example, Chrome DevTools).
+2. In the `/github-settings` dialog, inspect the list of labels.
+3. Locate one of the `
` elements representing a label.
+4. In the HTML, modify the `data-option-value` attribute of that `
` element
+ to the desired label name (for example, `release-failure`).
+5. Click on your modified label in the UI to select it, then save your settings.
diff --git a/docs/sidebar.json b/docs/sidebar.json
new file mode 100644
index 0000000000000000000000000000000000000000..bf82bc8220bfd04fb6c9acb86aeb08f8f7c33a1d
--- /dev/null
+++ b/docs/sidebar.json
@@ -0,0 +1,298 @@
+[
+ {
+ "label": "docs_tab",
+ "items": [
+ {
+ "label": "Get started",
+ "items": [
+ { "label": "Overview", "slug": "docs" },
+ { "label": "Quickstart", "slug": "docs/get-started" },
+ { "label": "Installation", "slug": "docs/get-started/installation" },
+ {
+ "label": "Authentication",
+ "slug": "docs/get-started/authentication"
+ },
+ { "label": "CLI cheatsheet", "slug": "docs/cli/cli-reference" },
+ {
+ "label": "Gemini 3 on Gemini CLI",
+ "slug": "docs/get-started/gemini-3"
+ }
+ ]
+ },
+ {
+ "label": "Use Gemini CLI",
+ "items": [
+ {
+ "label": "File management",
+ "slug": "docs/cli/tutorials/file-management"
+ },
+ {
+ "label": "Get started with Agent Skills",
+ "slug": "docs/cli/tutorials/skills-getting-started"
+ },
+ {
+ "label": "Manage context and memory",
+ "slug": "docs/cli/tutorials/memory-management"
+ },
+ {
+ "label": "Execute shell commands",
+ "slug": "docs/cli/tutorials/shell-commands"
+ },
+ {
+ "label": "Manage sessions and history",
+ "slug": "docs/cli/tutorials/session-management"
+ },
+ {
+ "label": "Plan tasks with todos",
+ "slug": "docs/cli/tutorials/task-planning"
+ },
+ {
+ "label": "Use Plan Mode with model steering",
+ "badge": "π¬",
+ "slug": "docs/cli/tutorials/plan-mode-steering"
+ },
+ {
+ "label": "Web search and fetch",
+ "slug": "docs/cli/tutorials/web-tools"
+ },
+ {
+ "label": "Set up an MCP server",
+ "slug": "docs/cli/tutorials/mcp-setup"
+ },
+ { "label": "Automate tasks", "slug": "docs/cli/tutorials/automation" }
+ ]
+ },
+ {
+ "label": "Features",
+ "items": [
+ {
+ "label": "Extensions",
+ "collapsed": true,
+ "items": [
+ {
+ "label": "Overview",
+ "slug": "docs/extensions"
+ },
+ {
+ "label": "User guide: Install and manage",
+ "link": "/docs/extensions/#manage-extensions"
+ },
+ {
+ "label": "Developer guide: Build extensions",
+ "slug": "docs/extensions/writing-extensions"
+ },
+ {
+ "label": "Developer guide: Best practices",
+ "slug": "docs/extensions/best-practices"
+ },
+ {
+ "label": "Developer guide: Releasing",
+ "slug": "docs/extensions/releasing"
+ },
+ {
+ "label": "Developer guide: Reference",
+ "slug": "docs/extensions/reference"
+ }
+ ]
+ },
+ {
+ "label": "Agent Skills",
+ "collapsed": true,
+ "items": [
+ { "label": "Overview", "slug": "docs/cli/skills" },
+ {
+ "label": "Get started with Agent Skills",
+ "slug": "docs/cli/tutorials/skills-getting-started"
+ },
+ {
+ "label": "Creating Agent Skills",
+ "slug": "docs/cli/creating-skills"
+ },
+ {
+ "label": "Using Agent Skills",
+ "slug": "docs/cli/using-agent-skills"
+ },
+ {
+ "label": "Developer guide: Best practices",
+ "slug": "docs/cli/skills-best-practices"
+ }
+ ]
+ },
+ {
+ "label": "Auto Memory",
+ "badge": "π¬",
+ "slug": "docs/cli/auto-memory"
+ },
+ { "label": "Checkpointing", "slug": "docs/cli/checkpointing" },
+ { "label": "Headless mode", "slug": "docs/cli/headless" },
+ {
+ "label": "Git worktrees",
+ "badge": "π¬",
+ "slug": "docs/cli/git-worktrees"
+ },
+ {
+ "label": "Hooks",
+ "collapsed": true,
+ "items": [
+ { "label": "Overview", "slug": "docs/hooks" },
+ { "label": "Reference", "slug": "docs/hooks/reference" }
+ ]
+ },
+ {
+ "label": "IDE integration",
+ "collapsed": true,
+ "items": [
+ { "label": "Overview", "slug": "docs/ide-integration" },
+ {
+ "label": "Developer guide: ACP mode",
+ "slug": "docs/cli/acp-mode"
+ }
+ ]
+ },
+ {
+ "label": "MCP servers",
+ "collapsed": true,
+ "items": [
+ { "label": "Overview", "slug": "docs/tools/mcp-server" },
+ { "label": "Resource tools", "slug": "docs/tools/mcp-resources" }
+ ]
+ },
+ { "label": "Model routing", "slug": "docs/cli/model-routing" },
+ { "label": "Model selection", "slug": "docs/cli/model" },
+ {
+ "label": "Model steering",
+ "badge": "π¬",
+ "slug": "docs/cli/model-steering"
+ },
+ {
+ "label": "Notifications",
+ "badge": "π¬",
+ "slug": "docs/cli/notifications"
+ },
+ { "label": "Plan mode", "slug": "docs/cli/plan-mode" },
+ {
+ "label": "Subagents",
+ "slug": "docs/core/subagents"
+ },
+ {
+ "label": "Remote subagents",
+ "slug": "docs/core/remote-agents"
+ },
+ { "label": "Rewind", "slug": "docs/cli/rewind" },
+ { "label": "Sandboxing", "slug": "docs/cli/sandbox" },
+ { "label": "Settings", "slug": "docs/cli/settings" },
+ { "label": "Telemetry", "slug": "docs/cli/telemetry" },
+ { "label": "Token caching", "slug": "docs/cli/token-caching" }
+ ]
+ },
+ {
+ "label": "Configuration",
+ "items": [
+ { "label": "Custom commands", "slug": "docs/cli/custom-commands" },
+ {
+ "label": "Enterprise configuration",
+ "slug": "docs/cli/enterprise"
+ },
+ {
+ "label": "Ignore files (.geminiignore)",
+ "slug": "docs/cli/gemini-ignore"
+ },
+ {
+ "label": "Model configuration",
+ "slug": "docs/cli/generation-settings"
+ },
+ {
+ "label": "Project context (GEMINI.md)",
+ "slug": "docs/cli/gemini-md"
+ },
+ { "label": "Settings", "slug": "docs/cli/settings" },
+ {
+ "label": "System prompt override",
+ "slug": "docs/cli/system-prompt"
+ },
+ { "label": "Themes", "slug": "docs/cli/themes" },
+ { "label": "Trusted folders", "slug": "docs/cli/trusted-folders" }
+ ]
+ },
+ {
+ "label": "Development",
+ "items": [
+ {
+ "label": "Behavioral evaluations",
+ "slug": "docs/behavioral-evals"
+ },
+ { "label": "Contribution guide", "slug": "docs/contributing" },
+ { "label": "Integration testing", "slug": "docs/integration-tests" },
+ {
+ "label": "Issue and PR automation",
+ "slug": "docs/issue-and-pr-automation"
+ },
+ { "label": "Local development", "slug": "docs/local-development" },
+ { "label": "NPM package structure", "slug": "docs/npm" }
+ ]
+ }
+ ]
+ },
+ {
+ "label": "reference_tab",
+ "items": [
+ {
+ "label": "Reference",
+ "items": [
+ { "label": "Command reference", "slug": "docs/reference/commands" },
+ {
+ "label": "Configuration reference",
+ "slug": "docs/reference/configuration"
+ },
+ {
+ "label": "Keyboard shortcuts",
+ "slug": "docs/reference/keyboard-shortcuts"
+ },
+ {
+ "label": "Memory import processor",
+ "slug": "docs/reference/memport"
+ },
+ { "label": "Policy engine", "slug": "docs/reference/policy-engine" },
+ { "label": "Tools reference", "slug": "docs/reference/tools" }
+ ]
+ }
+ ]
+ },
+ {
+ "label": "resources_tab",
+ "items": [
+ {
+ "label": "Resources",
+ "items": [
+ { "label": "FAQ", "slug": "docs/resources/faq" },
+ {
+ "label": "Quota and pricing",
+ "slug": "docs/resources/quota-and-pricing"
+ },
+ {
+ "label": "Terms and privacy",
+ "slug": "docs/resources/tos-privacy"
+ },
+ {
+ "label": "Troubleshooting",
+ "slug": "docs/resources/troubleshooting"
+ },
+ { "label": "Uninstall", "slug": "docs/resources/uninstall" }
+ ]
+ }
+ ]
+ },
+ {
+ "label": "releases_tab",
+ "items": [
+ {
+ "label": "Releases",
+ "items": [
+ { "label": "Release notes", "slug": "docs/changelogs/" },
+ { "label": "Stable release", "slug": "docs/changelogs/latest" },
+ { "label": "Preview release", "slug": "docs/changelogs/preview" }
+ ]
+ }
+ ]
+ }
+]
diff --git a/integration-tests/acp-env-auth.test.ts b/integration-tests/acp-env-auth.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..65f8adbf22a4290162ba30e0864ec2bb0cc895ca
--- /dev/null
+++ b/integration-tests/acp-env-auth.test.ts
@@ -0,0 +1,163 @@
+/**
+ * @license
+ * Copyright 2025 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, it, expect, beforeEach, afterEach } from 'vitest';
+import { TestRig } from './test-helper.js';
+import { spawn, ChildProcess } from 'node:child_process';
+import { join, resolve } from 'node:path';
+import { writeFileSync, mkdirSync } from 'node:fs';
+import { Writable, Readable } from 'node:stream';
+import { env } from 'node:process';
+import * as acp from '@agentclientprotocol/sdk';
+
+const sandboxEnv = env['GEMINI_SANDBOX'];
+const itMaybe = sandboxEnv && sandboxEnv !== 'false' ? it.skip : it;
+
+class MockClient implements acp.Client {
+ updates: acp.SessionNotification[] = [];
+ sessionUpdate = async (params: acp.SessionNotification) => {
+ this.updates.push(params);
+ };
+ requestPermission = async (): Promise => {
+ throw new Error('unexpected');
+ };
+}
+
+describe.skip('ACP Environment and Auth', () => {
+ let rig: TestRig;
+ let child: ChildProcess | undefined;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => {
+ child?.kill();
+ child = undefined;
+ await rig.cleanup();
+ });
+
+ itMaybe(
+ 'should load .env from project directory and use the provided API key',
+ async () => {
+ rig.setup('acp-env-loading');
+
+ // Create a project directory with a .env file containing a recognizable invalid key
+ const projectDir = resolve(join(rig.testDir!, 'project'));
+ mkdirSync(projectDir, { recursive: true });
+ writeFileSync(
+ join(projectDir, '.env'),
+ 'GEMINI_API_KEY=test-key-from-env\n',
+ );
+
+ const bundlePath = join(import.meta.dirname, '..', 'bundle/gemini.js');
+
+ child = spawn('node', [bundlePath, '--acp'], {
+ cwd: rig.homeDir!,
+ stdio: ['pipe', 'pipe', 'inherit'],
+ env: {
+ ...process.env,
+ GEMINI_CLI_HOME: rig.homeDir!,
+ GEMINI_API_KEY: undefined,
+ VERBOSE: 'true',
+ },
+ });
+
+ const input = Writable.toWeb(child.stdin!);
+ const output = Readable.toWeb(
+ child.stdout!,
+ ) as ReadableStream;
+ const testClient = new MockClient();
+ const stream = acp.ndJsonStream(input, output);
+ const connection = new acp.ClientSideConnection(() => testClient, stream);
+
+ await connection.initialize({
+ protocolVersion: acp.PROTOCOL_VERSION,
+ clientCapabilities: {
+ fs: { readTextFile: false, writeTextFile: false },
+ },
+ });
+
+ // 1. newSession should succeed because it finds the key in .env
+ const { sessionId } = await connection.newSession({
+ cwd: projectDir,
+ mcpServers: [],
+ });
+
+ expect(sessionId).toBeDefined();
+
+ // 2. prompt should fail because the key is invalid,
+ // but the error should come from the API, not the internal auth check.
+ await expect(
+ connection.prompt({
+ sessionId,
+ prompt: [{ type: 'text', text: 'hello' }],
+ }),
+ ).rejects.toSatisfy((error: unknown) => {
+ const acpError = error as acp.RequestError;
+ const errorData = acpError.data as
+ | { error?: { message?: string } }
+ | undefined;
+ const message = String(errorData?.error?.message || acpError.message);
+ // It should NOT be our internal "Authentication required" message
+ expect(message).not.toContain('Authentication required');
+ // It SHOULD be an API error mentioning the invalid key
+ expect(message).toContain('API key not valid');
+ return true;
+ });
+
+ child.stdin!.end();
+ },
+ );
+
+ itMaybe(
+ 'should fail with authRequired when no API key is found',
+ async () => {
+ rig.setup('acp-auth-failure');
+
+ const bundlePath = join(import.meta.dirname, '..', 'bundle/gemini.js');
+
+ child = spawn('node', [bundlePath, '--acp'], {
+ cwd: rig.homeDir!,
+ stdio: ['pipe', 'pipe', 'inherit'],
+ env: {
+ ...process.env,
+ GEMINI_CLI_HOME: rig.homeDir!,
+ GEMINI_API_KEY: undefined,
+ VERBOSE: 'true',
+ },
+ });
+
+ const input = Writable.toWeb(child.stdin!);
+ const output = Readable.toWeb(
+ child.stdout!,
+ ) as ReadableStream;
+ const testClient = new MockClient();
+ const stream = acp.ndJsonStream(input, output);
+ const connection = new acp.ClientSideConnection(() => testClient, stream);
+
+ await connection.initialize({
+ protocolVersion: acp.PROTOCOL_VERSION,
+ clientCapabilities: {
+ fs: { readTextFile: false, writeTextFile: false },
+ },
+ });
+
+ await expect(
+ connection.newSession({
+ cwd: resolve(rig.testDir!),
+ mcpServers: [],
+ }),
+ ).rejects.toMatchObject({
+ message: expect.stringContaining(
+ 'Gemini API key is missing or not configured.',
+ ),
+ });
+
+ child.stdin!.end();
+ },
+ );
+});
diff --git a/integration-tests/acp-telemetry.test.ts b/integration-tests/acp-telemetry.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..487dac474db6dd7584c3cdf0ecd0b36eabf90cb3
--- /dev/null
+++ b/integration-tests/acp-telemetry.test.ts
@@ -0,0 +1,116 @@
+/**
+ * @license
+ * Copyright 2025 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, it, expect, beforeEach, afterEach } from 'vitest';
+import { TestRig } from './test-helper.js';
+import { spawn, ChildProcess } from 'node:child_process';
+import { join } from 'node:path';
+import { readFileSync, existsSync } from 'node:fs';
+import { Writable, Readable } from 'node:stream';
+import { env } from 'node:process';
+import * as acp from '@agentclientprotocol/sdk';
+
+// Skip in sandbox mode - test spawns CLI directly which behaves differently in containers
+const sandboxEnv = env['GEMINI_SANDBOX'];
+const itMaybe = sandboxEnv && sandboxEnv !== 'false' ? it.skip : it;
+
+// Reuse existing fake responses that return a simple "Hello" response
+const SIMPLE_RESPONSE_PATH = 'hooks-system.session-startup.responses';
+
+class SessionUpdateCollector implements acp.Client {
+ updates: acp.SessionNotification[] = [];
+
+ sessionUpdate = async (params: acp.SessionNotification) => {
+ this.updates.push(params);
+ };
+
+ requestPermission = async (): Promise => {
+ throw new Error('unexpected');
+ };
+}
+
+describe('ACP telemetry', () => {
+ let rig: TestRig;
+ let child: ChildProcess | undefined;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => {
+ child?.kill();
+ child = undefined;
+ await rig.cleanup();
+ });
+
+ itMaybe('should flush telemetry when connection closes', async () => {
+ rig.setup('acp-telemetry-flush', {
+ fakeResponsesPath: join(import.meta.dirname, SIMPLE_RESPONSE_PATH),
+ });
+
+ const telemetryPath = join(rig.homeDir!, 'telemetry.log');
+ const bundlePath = join(import.meta.dirname, '..', 'bundle/gemini.js');
+
+ child = spawn(
+ 'node',
+ [
+ bundlePath,
+ '--acp',
+ '--fake-responses',
+ join(rig.testDir!, 'fake-responses.json'),
+ ],
+ {
+ cwd: rig.testDir!,
+ stdio: ['pipe', 'pipe', 'inherit'],
+ env: {
+ ...process.env,
+ GEMINI_API_KEY: 'fake-key',
+ GEMINI_CLI_HOME: rig.homeDir!,
+ GEMINI_TELEMETRY_ENABLED: 'true',
+ GEMINI_TELEMETRY_TRACES_ENABLED: 'true',
+ GEMINI_TELEMETRY_TARGET: 'local',
+ GEMINI_TELEMETRY_OUTFILE: telemetryPath,
+ },
+ },
+ );
+
+ const input = Writable.toWeb(child.stdin!);
+ const output = Readable.toWeb(child.stdout!) as ReadableStream;
+ const testClient = new SessionUpdateCollector();
+ const stream = acp.ndJsonStream(input, output);
+ const connection = new acp.ClientSideConnection(() => testClient, stream);
+
+ await connection.initialize({
+ protocolVersion: acp.PROTOCOL_VERSION,
+ clientCapabilities: { fs: { readTextFile: false, writeTextFile: false } },
+ });
+
+ const { sessionId } = await connection.newSession({
+ cwd: rig.testDir!,
+ mcpServers: [],
+ });
+
+ await connection.prompt({
+ sessionId,
+ prompt: [{ type: 'text', text: 'Say hello' }],
+ });
+
+ expect(JSON.stringify(testClient.updates)).toContain('Hello');
+
+ // Close stdin to trigger telemetry flush via runExitCleanup()
+ child.stdin!.end();
+ await new Promise((resolve) => {
+ child!.on('close', () => resolve());
+ });
+ child = undefined;
+
+ // gen_ai.output.messages is the last OTEL log emitted (after prompt response)
+ expect(existsSync(telemetryPath)).toBe(true);
+ expect(readFileSync(telemetryPath, 'utf-8')).toContain(
+ 'gen_ai.output.messages',
+ );
+ });
+});
diff --git a/integration-tests/api-resilience.responses b/integration-tests/api-resilience.responses
new file mode 100644
index 0000000000000000000000000000000000000000..d0520047f7e648e7216e14a9948e7cf268e53c79
--- /dev/null
+++ b/integration-tests/api-resilience.responses
@@ -0,0 +1 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Part 1. "}],"role":"model"},"index":0}]},{"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":10,"totalTokenCount":110}},{"candidates":[{"content":{"parts":[{"text":"Part 2."}],"role":"model"},"index":0,"finishReason":"STOP"}]}]}
diff --git a/integration-tests/api-resilience.test.ts b/integration-tests/api-resilience.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..870adf701a3063a371770e4275857a43193e0044
--- /dev/null
+++ b/integration-tests/api-resilience.test.ts
@@ -0,0 +1,50 @@
+/**
+ * @license
+ * Copyright 2026 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, it, expect, beforeEach, afterEach } from 'vitest';
+import { TestRig } from './test-helper.js';
+import { join, dirname } from 'node:path';
+import { fileURLToPath } from 'node:url';
+
+describe('API Resilience E2E', () => {
+ let rig: TestRig;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => {
+ await rig.cleanup();
+ });
+
+ it('should not crash when receiving metadata-only chunks in a stream', async () => {
+ await rig.setup('api-resilience-metadata-only', {
+ fakeResponsesPath: join(
+ dirname(fileURLToPath(import.meta.url)),
+ 'api-resilience.responses',
+ ),
+ settings: {
+ planSettings: { modelRouting: false },
+ },
+ });
+
+ // Run the CLI with a simple prompt.
+ // The fake responses will provide a stream with a metadata-only chunk in the middle.
+ // We use gemini-3-pro-preview to minimize internal service calls.
+ const result = await rig.run({
+ args: ['hi', '--model', 'gemini-3-pro-preview'],
+ });
+
+ // Verify the output contains text from the normal chunks.
+ // If the CLI crashed on the metadata chunk, rig.run would throw.
+ expect(result).toContain('Part 1.');
+ expect(result).toContain('Part 2.');
+
+ // Verify telemetry event for the prompt was still generated
+ const hasUserPromptEvent = await rig.waitForTelemetryEvent('user_prompt');
+ expect(hasUserPromptEvent).toBe(true);
+ });
+});
diff --git a/integration-tests/browser-agent-localhost.multistep.responses b/integration-tests/browser-agent-localhost.multistep.responses
new file mode 100644
index 0000000000000000000000000000000000000000..3ed786578f3c54547cc712885f5a009cf1e26c1c
--- /dev/null
+++ b/integration-tests/browser-agent-localhost.multistep.responses
@@ -0,0 +1,9 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll go through the multi-step flow on the localhost server."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Navigate to http://127.0.0.1:18923/multi-step/step1.html, fill in 'testuser' as the username, click Next, then on step 2 select 'Option B' and click Finish. Report the final result page content."}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"navigate_page","args":{"url":"http://127.0.0.1:18923/multi-step/step1.html"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":20,"totalTokenCount":120}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"fill","args":{"selector":"#username","value":"testuser"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":150,"candidatesTokenCount":25,"totalTokenCount":175}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"click","args":{"selector":"#next-btn"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":180,"candidatesTokenCount":20,"totalTokenCount":200}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":210,"candidatesTokenCount":15,"totalTokenCount":225}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"click","args":{"selector":"#finish-btn"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":240,"candidatesTokenCount":20,"totalTokenCount":260}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":270,"candidatesTokenCount":15,"totalTokenCount":285}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"result":{"success":true,"summary":"Completed all steps. Step 1: entered username 'testuser'. Step 2: selected default option. Final result page shows 'Multi-Step Complete' with 'β Complete' status badge."}}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":300,"candidatesTokenCount":40,"totalTokenCount":340}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I've completed the multi-step flow:\n\n1. **Step 1**: Entered 'testuser' as username and clicked Next\n2. **Step 2**: Confirmed selection and clicked Finish\n3. **Result**: Final page shows 'Multi-Step Complete' with a 'β Complete' status badge\n\nAll steps were successfully navigated."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":300,"candidatesTokenCount":60,"totalTokenCount":360}}]}
diff --git a/integration-tests/browser-agent-localhost.navigate.responses b/integration-tests/browser-agent-localhost.navigate.responses
new file mode 100644
index 0000000000000000000000000000000000000000..7c25e8294570fbeaf7d0a83685c09b433ab06f27
--- /dev/null
+++ b/integration-tests/browser-agent-localhost.navigate.responses
@@ -0,0 +1,5 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll navigate to the localhost page and read its content using the browser agent."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Navigate to http://127.0.0.1:18923/index.html and tell me the page title and list all links on the page"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":40,"totalTokenCount":140}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"navigate_page","args":{"url":"http://127.0.0.1:18923/index.html"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":20,"totalTokenCount":120}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":150,"candidatesTokenCount":20,"totalTokenCount":170}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"result":{"success":true,"summary":"Page title is 'Test Fixture - Home'. Found 3 links: Contact Form (/form.html), Multi-Step Flow (/multi-step/step1.html), Dynamic Content (/dynamic.html)."}}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":40,"totalTokenCount":240}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"The localhost test fixture page has:\n\n**Title**: Test Fixture - Home\n\n**Links**:\n1. Contact Form (form.html)\n2. Multi-Step Flow (multi-step/step1.html)\n3. Dynamic Content (dynamic.html)\n\nThe page also has a heading 'Test Fixture Home Page' and footer content."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":60,"totalTokenCount":260}}]}
diff --git a/integration-tests/browser-agent.cleanup.responses b/integration-tests/browser-agent.cleanup.responses
new file mode 100644
index 0000000000000000000000000000000000000000..755341ef0f9e842dba2293580116c1de8fda92b4
--- /dev/null
+++ b/integration-tests/browser-agent.cleanup.responses
@@ -0,0 +1,5 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll open https://example.com and check the page title for you."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Open https://example.com and get the page title"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":35,"totalTokenCount":135}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"navigate_page","args":{"url":"https://example.com"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":20,"totalTokenCount":120}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":150,"candidatesTokenCount":20,"totalTokenCount":170}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"result":{"success":true,"summary":"The page title is 'Example Domain'."}}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":30,"totalTokenCount":230}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I have opened the page and the title is 'Example Domain'. The browser session has been cleaned up successfully."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":30,"totalTokenCount":230}}]}
diff --git a/integration-tests/browser-agent.confirmation.responses b/integration-tests/browser-agent.confirmation.responses
new file mode 100644
index 0000000000000000000000000000000000000000..4f645c6531ff27e46d53d43619d89ae458a69a27
--- /dev/null
+++ b/integration-tests/browser-agent.confirmation.responses
@@ -0,0 +1 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"file_path":"test.txt","content":"hello"}}},{"text":"I've successfully written \"hello\" to test.txt. The file has been created with the specified content."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
diff --git a/integration-tests/browser-policy.responses b/integration-tests/browser-policy.responses
new file mode 100644
index 0000000000000000000000000000000000000000..95b055d5c716ed2fb91697f1b3c75d01d0133e7b
--- /dev/null
+++ b/integration-tests/browser-policy.responses
@@ -0,0 +1,5 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll help you with that."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Open https://example.com and check if there is a heading"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"new_page","args":{"url":"https://example.com"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"success":true,"summary":"SUCCESS_POLICY_TEST_COMPLETED"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Task completed successfully. The page has the heading \"Example Domain\"."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":50,"totalTokenCount":250}}]}
diff --git a/integration-tests/browser-policy.test.ts b/integration-tests/browser-policy.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..d727ca2fc126f2381396e14116e8d3374b1f42c2
--- /dev/null
+++ b/integration-tests/browser-policy.test.ts
@@ -0,0 +1,240 @@
+/**
+ * @license
+ * Copyright 2026 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, it, expect, beforeEach, afterEach } from 'vitest';
+import { TestRig, poll } from './test-helper.js';
+import { dirname, join } from 'node:path';
+import { fileURLToPath } from 'node:url';
+import { execSync } from 'node:child_process';
+import { existsSync, writeFileSync, readFileSync, mkdirSync } from 'node:fs';
+import { env } from 'node:process';
+import stripAnsi from 'strip-ansi';
+
+// Browser agent Chrome DevTools MCP connection is flaky in Docker sandbox.
+// See: https://github.com/google-gemini/gemini-cli/issues/24382
+const isDockerSandbox = env['GEMINI_SANDBOX'] === 'docker';
+
+const __filename = fileURLToPath(import.meta.url);
+const __dirname = dirname(__filename);
+
+const chromeAvailable = (() => {
+ try {
+ if (process.platform === 'darwin') {
+ execSync(
+ 'test -d "/Applications/Google Chrome.app" || test -d "/Applications/Chromium.app"',
+ {
+ stdio: 'ignore',
+ },
+ );
+ } else if (process.platform === 'linux') {
+ execSync(
+ 'which google-chrome || which chromium-browser || which chromium',
+ { stdio: 'ignore' },
+ );
+ } else if (process.platform === 'win32') {
+ const chromePaths = [
+ 'C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe',
+ 'C:\\Program Files (x86)\\Google\\Chrome\\Application\\chrome.exe',
+ `${process.env['LOCALAPPDATA'] ?? ''}\\Google\\Chrome\\Application\\chrome.exe`,
+ ];
+ const found = chromePaths.some((p) => existsSync(p));
+ if (!found) {
+ execSync('where chrome || where chromium', { stdio: 'ignore' });
+ }
+ } else {
+ return false;
+ }
+ return true;
+ } catch {
+ return false;
+ }
+})();
+
+describe.skipIf(!chromeAvailable)('browser-policy', () => {
+ let rig: TestRig;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => {
+ await rig.cleanup();
+ });
+
+ it.skipIf(isDockerSandbox)(
+ 'should skip confirmation when "Allow all server tools for this session" is chosen',
+ async () => {
+ rig.setup('browser-policy-skip-confirmation', {
+ fakeResponsesPath: join(__dirname, 'browser-policy.responses'),
+ settings: {
+ agents: {
+ overrides: {
+ browser_agent: {
+ enabled: true,
+ },
+ },
+ browser: {
+ headless: true,
+ sessionMode: 'isolated',
+ allowedDomains: ['example.com'],
+ },
+ },
+ },
+ });
+
+ // Manually trust the folder to avoid the dialog and enable option 3
+ const geminiDir = join(rig.homeDir!, '.gemini');
+ mkdirSync(geminiDir, { recursive: true });
+
+ // Write to trustedFolders.json
+ const trustedFoldersPath = join(geminiDir, 'trustedFolders.json');
+ const trustedFolders = {
+ [rig.testDir!]: 'TRUST_FOLDER',
+ };
+ writeFileSync(
+ trustedFoldersPath,
+ JSON.stringify(trustedFolders, null, 2),
+ );
+
+ // Force confirmation for browser agent.
+ // NOTE: We don't force confirm browser tools here because "Allow all server tools"
+ // adds a rule with ALWAYS_ALLOW_PRIORITY (3.9x) which would be overshadowed by
+ // a rule in the user tier (4.x) like the one from this TOML.
+ // By removing the explicit mcp rule, the first MCP tool will still prompt
+ // due to default approvalMode = 'default', and then "Allow all" will correctly
+ // bypass subsequent tools.
+ const policyFile = join(rig.testDir!, 'force-confirm.toml');
+ writeFileSync(
+ policyFile,
+ `
+[[rule]]
+name = "Force confirm browser_agent"
+toolName = "invoke_agent"
+argsPattern = "\\"agent_name\\":\\\\s*\\"browser_agent\\""
+decision = "ask_user"
+priority = 200
+`,
+ );
+
+ // Update settings.json in both project and home directories to point to the policy file
+ for (const baseDir of [rig.testDir!, rig.homeDir!]) {
+ const settingsPath = join(baseDir, '.gemini', 'settings.json');
+ if (existsSync(settingsPath)) {
+ const settings = JSON.parse(readFileSync(settingsPath, 'utf-8'));
+ settings.policyPaths = [policyFile];
+ // Ensure folder trust is enabled
+ settings.security = settings.security || {};
+ settings.security.folderTrust = settings.security.folderTrust || {};
+ settings.security.folderTrust.enabled = true;
+ writeFileSync(settingsPath, JSON.stringify(settings, null, 2));
+ }
+ }
+
+ const run = await rig.runInteractive({
+ approvalMode: 'default',
+ env: {
+ GEMINI_CLI_INTEGRATION_TEST: 'true',
+ },
+ });
+
+ await run.sendKeys(
+ 'Open https://example.com and check if there is a heading\r',
+ );
+ await run.sendKeys('\r');
+
+ // Handle confirmations.
+ // 1. Initial browser_agent delegation (likely only 3 options, so use option 1: Allow once)
+ await poll(
+ () => stripAnsi(run.output).toLowerCase().includes('action required'),
+ 60000,
+ 1000,
+ );
+ await run.sendKeys('1\r');
+ await new Promise((r) => setTimeout(r, 2000));
+
+ // Handle privacy notice
+ await poll(
+ () => stripAnsi(run.output).toLowerCase().includes('privacy notice'),
+ 5000,
+ 100,
+ );
+ await run.sendKeys('1\r');
+ await new Promise((r) => setTimeout(r, 5000));
+
+ // new_page (MCP tool, should have 4 options, use option 3: Allow all server tools)
+ await poll(
+ () => {
+ const stripped = stripAnsi(run.output).toLowerCase();
+ return (
+ stripped.includes('new_page') &&
+ stripped.includes('allow all server tools for this session')
+ );
+ },
+ 60000,
+ 1000,
+ );
+
+ // Select "Allow all server tools for this session" (option 3)
+ await run.sendKeys('3\r');
+
+ // Wait for the browser agent to finish (success or failure)
+ await poll(
+ () => {
+ const stripped = stripAnsi(run.output).toLowerCase();
+ return (
+ stripped.includes('completed successfully') ||
+ stripped.includes('agent error')
+ );
+ },
+ 120000,
+ 1000,
+ );
+
+ const output = stripAnsi(run.output).toLowerCase();
+
+ expect(output).toContain('browser_agent');
+ // The test validates that "Allow all server tools" skips subsequent
+ // tool confirmations β the browser agent may still fail due to
+ // Chrome/MCP issues in CI, which is acceptable for this policy test.
+ expect(
+ output.includes('completed successfully') ||
+ output.includes('agent error'),
+ ).toBe(true);
+ },
+ );
+
+ it('should show the visible warning when browser agent starts in existing session mode', async () => {
+ rig.setup('browser-session-warning', {
+ fakeResponsesPath: join(__dirname, 'browser-agent.cleanup.responses'),
+ settings: {
+ general: {
+ enableAutoUpdateNotification: false,
+ },
+ agents: {
+ overrides: {
+ browser_agent: {
+ enabled: true,
+ },
+ },
+ browser: {
+ sessionMode: 'existing',
+ headless: true,
+ },
+ },
+ },
+ });
+
+ const stdout = await rig.runCommand(['Open https://example.com'], {
+ env: {
+ GEMINI_API_KEY: 'fake-key',
+ GEMINI_TELEMETRY_DISABLED: 'true',
+ DEV: 'true',
+ },
+ });
+
+ expect(stdout).toContain('saved logins will be visible');
+ });
+});
diff --git a/integration-tests/checkpointing.test.ts b/integration-tests/checkpointing.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..72277f25dafc8415d31afae03932b419528f23d7
--- /dev/null
+++ b/integration-tests/checkpointing.test.ts
@@ -0,0 +1,155 @@
+/**
+ * @license
+ * Copyright 2025 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, it, expect, beforeEach, afterEach } from 'vitest';
+import * as fs from 'node:fs/promises';
+import * as path from 'node:path';
+import * as os from 'node:os';
+import { GitService, Storage } from '@google/gemini-cli-core';
+
+describe('Checkpointing Integration', () => {
+ let tmpDir: string;
+ let projectRoot: string;
+ let fakeHome: string;
+ let originalEnv: NodeJS.ProcessEnv;
+
+ beforeEach(async () => {
+ tmpDir = await fs.mkdtemp(
+ path.join(os.tmpdir(), 'gemini-checkpoint-test-'),
+ );
+ projectRoot = path.join(tmpDir, 'project');
+ fakeHome = path.join(tmpDir, 'home');
+
+ await fs.mkdir(projectRoot, { recursive: true });
+ await fs.mkdir(fakeHome, { recursive: true });
+
+ // Save original env
+ originalEnv = { ...process.env };
+
+ // Simulate environment with NO global gitconfig
+ process.env['HOME'] = fakeHome;
+ delete process.env['GIT_CONFIG_GLOBAL'];
+ delete process.env['GIT_CONFIG_SYSTEM'];
+ });
+
+ afterEach(async () => {
+ // Restore env
+ process.env = originalEnv;
+
+ // Cleanup
+ try {
+ await fs.rm(tmpDir, { recursive: true, force: true });
+ } catch (e) {
+ console.error('Failed to cleanup temp dir', e);
+ }
+ });
+
+ it('should successfully create and restore snapshots without global git config', async () => {
+ const storage = new Storage(projectRoot);
+ const gitService = new GitService(projectRoot, storage);
+
+ // 1. Initialize
+ await gitService.initialize();
+
+ // Verify system config empty file creation
+ // We need to access getHistoryDir logic or replicate it.
+ // Since we don't have access to private getHistoryDir, we can infer it or just trust the functional test.
+
+ // 2. Create initial state
+ await fs.writeFile(path.join(projectRoot, 'file1.txt'), 'version 1');
+ await fs.writeFile(path.join(projectRoot, 'file2.txt'), 'permanent file');
+
+ // 3. Create Snapshot
+ const snapshotHash = await gitService.createFileSnapshot('Checkpoint 1');
+ expect(snapshotHash).toBeDefined();
+
+ // 4. Modify files
+ await fs.writeFile(
+ path.join(projectRoot, 'file1.txt'),
+ 'version 2 (BAD CHANGE)',
+ );
+ await fs.writeFile(
+ path.join(projectRoot, 'file3.txt'),
+ 'new file (SHOULD BE GONE)',
+ );
+ await fs.rm(path.join(projectRoot, 'file2.txt'));
+
+ // 5. Restore
+ await gitService.restoreProjectFromSnapshot(snapshotHash);
+
+ // 6. Verify state
+ const file1Content = await fs.readFile(
+ path.join(projectRoot, 'file1.txt'),
+ 'utf-8',
+ );
+ expect(file1Content).toBe('version 1');
+
+ const file2Exists = await fs
+ .stat(path.join(projectRoot, 'file2.txt'))
+ .then(() => true)
+ .catch(() => false);
+ expect(file2Exists).toBe(true);
+ const file2Content = await fs.readFile(
+ path.join(projectRoot, 'file2.txt'),
+ 'utf-8',
+ );
+ expect(file2Content).toBe('permanent file');
+
+ const file3Exists = await fs
+ .stat(path.join(projectRoot, 'file3.txt'))
+ .then(() => true)
+ .catch(() => false);
+ expect(file3Exists).toBe(false);
+ });
+
+ it('should ignore user global git config and use isolated identity', async () => {
+ // 1. Create a fake global gitconfig with a specific user
+ const globalConfigPath = path.join(fakeHome, '.gitconfig');
+ const globalConfigContent = `[user]
+ name = Global User
+ email = global@example.com
+`;
+ await fs.writeFile(globalConfigPath, globalConfigContent);
+
+ // Point HOME to fakeHome so git picks up this global config (if we didn't isolate it)
+ process.env['HOME'] = fakeHome;
+ // Ensure GIT_CONFIG_GLOBAL is NOT set for the process initially,
+ // so it would default to HOME/.gitconfig if GitService didn't override it.
+ delete process.env['GIT_CONFIG_GLOBAL'];
+
+ const storage = new Storage(projectRoot);
+ const gitService = new GitService(projectRoot, storage);
+
+ await gitService.initialize();
+
+ // 2. Create a file and snapshot
+ await fs.writeFile(path.join(projectRoot, 'test.txt'), 'content');
+ await gitService.createFileSnapshot('Snapshot with global config present');
+
+ // 3. Verify the commit author in the shadow repo
+ const historyDir = storage.getHistoryDir();
+
+ const { execFileSync } = await import('node:child_process');
+
+ const logOutput = execFileSync(
+ 'git',
+ ['log', '-1', '--pretty=format:%an <%ae>'],
+ {
+ cwd: historyDir,
+ env: {
+ ...process.env,
+ GIT_DIR: path.join(historyDir, '.git'),
+ GIT_CONFIG_GLOBAL: path.join(historyDir, '.gitconfig'),
+ GIT_CONFIG_SYSTEM: path.join(historyDir, '.gitconfig_system_empty'),
+ },
+ encoding: 'utf-8',
+ },
+ );
+
+ expect(logOutput).toBe('Gemini CLI ');
+ expect(logOutput).not.toContain('Global User');
+ });
+});
diff --git a/integration-tests/concurrency-limit.responses b/integration-tests/concurrency-limit.responses
new file mode 100644
index 0000000000000000000000000000000000000000..e2bd5efe2aefb3735df57b65108f9e5cc1796416
--- /dev/null
+++ b/integration-tests/concurrency-limit.responses
@@ -0,0 +1,12 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/1"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/2"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/3"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/4"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/5"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/6"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/7"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/8"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/9"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/10"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/11"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":500,"totalTokenCount":600}}]}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 1 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 2 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 3 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 4 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 5 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 6 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 7 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 8 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 9 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 10 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Some requests were rate limited: Rate limit exceeded for host. Please wait 60 seconds before trying again."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":1000,"candidatesTokenCount":50,"totalTokenCount":1050}}]}
diff --git a/integration-tests/context-compress-interactive.compress-empty.responses b/integration-tests/context-compress-interactive.compress-empty.responses
new file mode 100644
index 0000000000000000000000000000000000000000..e69de29bb2d1d6434b8b29ae775ad8c2e48c5391
diff --git a/integration-tests/context-compress-interactive.compress-failure.responses b/integration-tests/context-compress-interactive.compress-failure.responses
new file mode 100644
index 0000000000000000000000000000000000000000..7ba10591a6251a62385433846efd1f6d25624741
--- /dev/null
+++ b/integration-tests/context-compress-interactive.compress-failure.responses
@@ -0,0 +1,3 @@
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"thought":true,"text":"**Observing Initial Conditions**\n\nI'm currently focused on the initial context. I've taken note of the provided date, OS, and working directory. I'm also carefully examining the file structure presented within the current working directory. It's helping me understand the starting point for further analysis.\n\n\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12270,"totalTokenCount":12316,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12270}],"thoughtsTokenCount":46}},{"candidates":[{"content":{"parts":[{"thought":true,"text":"**Assessing User Intent**\n\nI'm now shifting my focus. I've successfully registered the provided data and file structure. My current task is to understand the user's ultimate goal, given the information provided. The \"Hello.\" command is straightforward, but I'm checking if there's an underlying objective.\n\n\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12270,"totalTokenCount":12341,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12270}],"thoughtsTokenCount":71}},{"candidates":[{"content":{"parts":[{"thoughtSignature":"CiQB0e2Kb3dRh+BYdbZvmulSN2Pwbc75DfQOT3H4EN0rn039hoMKfwHR7YpvvyqNKoxXAiCbYw3gbcTr/+pegUpgnsIrt8oQPMytFMjKSsMyshfygc21T2MkyuI6Q5I/fNCcHROWexdZnIeppVCDB2TarN4LGW4T9Yci6n/ynMMFT2xc2/vyHpkDgRM7avhMElnBhuxAY+e4TpxkZIncGWCEHP1TouoKpgEB0e2Kb8Xpwm0hiKhPt2ZLizpxjk+CVtcbnlgv69xo5VsuQ+iNyrVGBGRwNx+eTeNGdGpn6e73WOCZeP91FwOZe7URyL12IA6E6gYWqw0kXJR4hO4p6Lwv49E3+FRiG2C4OKDF8LF5XorYyCHSgBFT1/RUAVj81GDTx1xxtmYKN3xq8Ri+HsPbqU/FM/jtNZKkXXAtufw2Bmw8lJfmugENIv/TQI7xCo8BAdHtim8KgAXJfZ7ASfutVLKTylQeaslyB/SmcHJ0ZiNr5j8WP1prZdb6XnZZ1ZNbhjxUf/ymoxHKGvtTPBgLE9azMj8Lx/k0clhd2a+wNsiIqW9qCzlVah0tBMytpQUjIDtQe9Hj4LLUprF9PUe/xJkj000Z0ZzsgFm2ncdTWZTdkhCQDpyETVAxdE+oklwKJAHR7YpvUjSkD6KwY1gLrOsHKy0UNfn2lMbxjVetKNMVBRqsTg==","text":"Hello."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12270,"totalTokenCount":12341,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12270}],"thoughtsTokenCount":71}}]}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"\n \n \n \n\n \n - OS: linux\n - Date: Friday, October 24, 2025\n \n\n \n - OBSERVED: The directory contains `telemetry.log` and a `.gemini/` directory.\n - OBSERVED: The `.gemini/` directory contains `settings.json` and `settings.json.orig`.\n \n\n \n - The user initiated the chat.\n \n\n \n 1. [TODO] Await the user's first instruction to formulate a plan.\n \n"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":983,"candidatesTokenCount":299,"totalTokenCount":1637,"promptTokensDetails":[{"modality":"TEXT","tokenCount":983}],"thoughtsTokenCount":355}}}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"\n \n \n \n\n \n - OS: linux\n - Date: Friday, October 24, 2025\n \n\n \n - OBSERVED: The directory contains `telemetry.log` and a `.gemini/` directory.\n - OBSERVED: The `.gemini/` directory contains `settings.json` and `settings.json.orig`.\n \n\n \n - The user initiated the chat.\n \n\n \n 1. [TODO] Await the user's first instruction to formulate a plan.\n \n"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":983,"candidatesTokenCount":299,"totalTokenCount":1637,"promptTokensDetails":[{"modality":"TEXT","tokenCount":983}],"thoughtsTokenCount":355}}}
diff --git a/integration-tests/context-fidelity.test.ts b/integration-tests/context-fidelity.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..845b25b22f42c16caade16af9031a073f0e9de0e
--- /dev/null
+++ b/integration-tests/context-fidelity.test.ts
@@ -0,0 +1,287 @@
+/**
+ * @license
+ * Copyright 2026 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, it, expect, beforeEach, afterEach } from 'vitest';
+import { TestRig } from './test-helper.js';
+import * as path from 'node:path';
+import * as fs from 'node:fs';
+import { FinishReason, GenerateContentResponse } from '@google/genai';
+import type { FakeResponse, HistoryTurn } from '@google/gemini-cli-core';
+
+describe('Context Management Fidelity E2E', () => {
+ let rig: TestRig;
+
+ function generateRandomString(length: number): string {
+ const characters =
+ 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789';
+ let result = '';
+ for (let i = 0; i < length; i++) {
+ result += characters.charAt(
+ Math.floor(Math.random() * characters.length),
+ );
+ }
+ return result;
+ }
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => await rig.cleanup());
+
+ it(
+ 'should reproduce the exact context working buffer on resume',
+ { timeout: 300000 },
+ async () => {
+ // Mock responses to trigger GC (summarization)
+ const snapshotResponse: FakeResponse = {
+ method: 'generateContent',
+ response: {
+ candidates: [
+ {
+ content: {
+ parts: [
+ {
+ text: JSON.stringify({
+ new_facts: ['GC Triggered.'],
+ new_constraints: [],
+ new_tasks: [],
+ resolved_task_ids: [],
+ obsolete_fact_indices: [],
+ obsolete_constraint_indices: [],
+ chronological_summary: 'Snapshot created.',
+ }),
+ },
+ ],
+ role: 'model',
+ },
+ finishReason: FinishReason.STOP,
+ index: 0,
+ },
+ ],
+ } as unknown as GenerateContentResponse,
+ };
+
+ const countTokensResponse: FakeResponse = {
+ method: 'countTokens',
+ response: { totalTokens: 1000 },
+ };
+
+ const streamResponse = (text: string): FakeResponse => ({
+ method: 'generateContentStream',
+ response: [
+ {
+ candidates: [
+ {
+ content: { parts: [{ text }], role: 'model' },
+ finishReason: FinishReason.STOP,
+ index: 0,
+ },
+ ],
+ },
+ ] as unknown as GenerateContentResponse[],
+ });
+
+ const setupResponses = (fileName: string, mocks: FakeResponse[]) => {
+ const filePath = path.join(rig.testDir!, fileName);
+ fs.writeFileSync(
+ filePath,
+ mocks.map((m) => JSON.stringify(m)).join('\n'),
+ );
+ return filePath;
+ };
+
+ await rig.setup('context-fidelity', {
+ settings: {
+ experimental: {
+ stressTestProfile: true, // Lowers thresholds to trigger GC easily
+ },
+ },
+ });
+
+ const traceDir = path.join(rig.testDir!, 'traces');
+ fs.mkdirSync(traceDir, { recursive: true });
+ const traceLog = path.join(traceDir, 'trace.log');
+
+ // Ignore trace and response files to keep environment context clean and stable
+ fs.writeFileSync(
+ path.join(rig.testDir!, '.geminiignore'),
+ 'traces/\nresp*.json\ndebug.log\n',
+ );
+
+ const commonEnv = {
+ GEMINI_API_KEY: 'mock-key',
+ GEMINI_CONTEXT_TRACE_DIR: traceDir,
+ GEMINI_CONTEXT_TRACE_ENABLED: 'true',
+ GEMINI_DEBUG_LOG_FILE: path.join(rig.testDir!, 'debug.log'),
+ };
+
+ const runMocks: FakeResponse[] = [
+ streamResponse('Ack 1'),
+ streamResponse('Ack 2'),
+ streamResponse('Ack 3'),
+ streamResponse('Ack 4'),
+ streamResponse('Ack 5'),
+ streamResponse('Ack 6'),
+ streamResponse('Ack 7'),
+ streamResponse('Ack 8'),
+ streamResponse('Ack 9'),
+ streamResponse('Ack 10'),
+ streamResponse('Ack 11'),
+ streamResponse('Ack 12'),
+ ];
+ for (let i = 0; i < 50; i++) {
+ runMocks.push(snapshotResponse);
+ runMocks.push(countTokensResponse);
+ }
+
+ // Turns 1-10: Build up history
+ for (let i = 1; i <= 10; i++) {
+ await rig.run({
+ args: [
+ '--debug',
+ i === 1 ? '' : '--resume',
+ i === 1 ? '' : 'latest',
+ '--fake-responses-non-strict',
+ setupResponses(`resp_init_${i}.json`, runMocks),
+ ].filter(Boolean),
+ stdin: `Turn ${i}: ` + generateRandomString(900),
+ env: commonEnv,
+ });
+ }
+
+ // Turn 11: Penultimate turn
+ await rig.run({
+ args: [
+ '--debug',
+ '--resume',
+ 'latest',
+ '--fake-responses-non-strict',
+ setupResponses('resp2.json', runMocks),
+ ],
+ stdin: 'Turn 11: ' + generateRandomString(900),
+ env: commonEnv,
+ });
+
+ // Turn 12: Breach threshold and force GC
+ await rig.run({
+ args: [
+ '--debug',
+ '--resume',
+ 'latest',
+ '--fake-responses-non-strict',
+ setupResponses('resp3.json', runMocks),
+ ],
+ stdin: 'Turn 12: ' + generateRandomString(900),
+ env: commonEnv,
+ });
+
+ // Extract the rendered context asset from the log
+ const getRenderedContext = (logContent: string): HistoryTurn[] | null => {
+ const lines = logContent.split('\n');
+ const renderLines = lines.filter(
+ (l) =>
+ l.includes('[Render] Render Sanitized Context for LLM') ||
+ l.includes('[Render] Render Context for LLM'),
+ );
+ if (renderLines.length === 0) return null;
+
+ const lastRender = renderLines[renderLines.length - 1];
+ const detailsMatch = lastRender.match(/\| Details: (.*)$/);
+ if (!detailsMatch) return null;
+
+ const details = JSON.parse(detailsMatch[1]);
+ const assetInfo =
+ details.renderedContextSanitized || details.renderedContext;
+ if (assetInfo && assetInfo.$asset) {
+ const assetPath = path.join(traceDir, 'assets', assetInfo.$asset);
+ return JSON.parse(fs.readFileSync(assetPath, 'utf-8'));
+ }
+ return assetInfo;
+ };
+
+ const log1 = fs.readFileSync(traceLog, 'utf-8');
+ const contextBeforeExit = getRenderedContext(log1);
+ expect(contextBeforeExit).toBeDefined();
+ console.log(
+ 'Context Before Exit (First 2 turns):',
+ JSON.stringify(contextBeforeExit!.slice(0, 2), null, 2),
+ );
+
+ // Turn 4: Resume and run a small command
+ await rig.run({
+ args: [
+ '--debug',
+ '--resume',
+ 'latest',
+ '--fake-responses-non-strict',
+ setupResponses('resp4.json', runMocks),
+ 'continue',
+ ],
+ env: commonEnv,
+ });
+
+ const log2 = fs.readFileSync(traceLog, 'utf-8');
+ const contextAfterResume = getRenderedContext(log2);
+ expect(contextAfterResume).toBeDefined();
+ console.log(
+ 'Context After Resume (First 2 turns):',
+ JSON.stringify(contextAfterResume!.slice(0, 2), null, 2),
+ );
+
+ expect(contextAfterResume!.length).toBeGreaterThanOrEqual(
+ contextBeforeExit!.length,
+ );
+
+ // The environment context is intentionally refreshed on resume to reflect
+ // the current state of the workspace (e.g. new files, current date).
+ // We allow its content to differ but ensure it's still an environment context.
+ const isEnvContext = (turn: HistoryTurn) =>
+ turn.content.parts?.some((p) => p.text?.includes(''));
+
+ for (let i = 0; i < contextBeforeExit!.length; i++) {
+ expect(contextAfterResume![i].id).toBe(contextBeforeExit![i].id);
+
+ const turnBefore = contextBeforeExit![i];
+ const turnAfter = contextAfterResume![i];
+
+ if (isEnvContext(turnBefore)) {
+ expect(isEnvContext(turnAfter)).toBe(true);
+ continue;
+ }
+
+ expect(turnAfter.content).toEqual(turnBefore.content);
+ }
+
+ // Most importantly, synthetic IDs (like summaries) must be stable.
+ const syntheticTurns = contextBeforeExit!.filter(
+ (t: HistoryTurn) =>
+ t.content.parts?.some((p) => p.text?.includes('active_tasks')) ||
+ (t.id && t.id.length === 32),
+ );
+ expect(syntheticTurns.length).toBeGreaterThan(0);
+
+ const syntheticTurnsAfter = contextAfterResume!.filter(
+ (t: HistoryTurn) =>
+ t.content.parts?.some((p) => p.text?.includes('active_tasks')) ||
+ (t.id && t.id.length === 32),
+ );
+ expect(syntheticTurnsAfter.length).toBeGreaterThanOrEqual(
+ syntheticTurns.length,
+ );
+
+ // Check if the first synthetic turn is identical (with relaxation for environment context)
+ expect(syntheticTurnsAfter[0].id).toBe(syntheticTurns[0].id);
+ if (isEnvContext(syntheticTurns[0])) {
+ expect(isEnvContext(syntheticTurnsAfter[0])).toBe(true);
+ } else {
+ expect(syntheticTurnsAfter[0].content).toEqual(
+ syntheticTurns[0].content,
+ );
+ }
+ },
+ );
+});
diff --git a/integration-tests/extensions-install.test.ts b/integration-tests/extensions-install.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..e9f1cdbf49ea1ccb9df6be5584c0af07477d7d5d
--- /dev/null
+++ b/integration-tests/extensions-install.test.ts
@@ -0,0 +1,62 @@
+/**
+ * @license
+ * Copyright 2025 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { describe, expect, it, beforeEach, afterEach } from 'vitest';
+import { TestRig } from './test-helper.js';
+import { writeFileSync } from 'node:fs';
+import { join } from 'node:path';
+
+const extension = `{
+ "name": "test-extension-install",
+ "version": "0.0.1"
+}`;
+
+const extensionUpdate = `{
+ "name": "test-extension-install",
+ "version": "0.0.2"
+}`;
+
+describe('extension install', () => {
+ let rig: TestRig;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => await rig.cleanup());
+
+ it('installs a local extension, verifies a command, and updates it', async () => {
+ rig.setup('extension install test');
+ const testServerPath = join(rig.testDir!, 'gemini-extension.json');
+ writeFileSync(testServerPath, extension);
+ try {
+ const result = await rig.runCommand(
+ ['--debug', 'extensions', 'install', `${rig.testDir!}`],
+ { stdin: 'y\n' },
+ );
+ expect(result).toContain('test-extension-install');
+
+ const listResult = await rig.runCommand([
+ '--debug',
+ 'extensions',
+ 'list',
+ ]);
+ expect(listResult).toContain('test-extension-install');
+ writeFileSync(testServerPath, extensionUpdate);
+ const updateResult = await rig.runCommand(
+ ['--debug', 'extensions', 'update', `test-extension-install`],
+ { stdin: 'y\n' },
+ );
+ expect(updateResult).toContain('0.0.2');
+ } finally {
+ await rig.runCommand([
+ 'extensions',
+ 'uninstall',
+ 'test-extension-install',
+ ]);
+ }
+ });
+});
diff --git a/integration-tests/extensions-reload.test.ts b/integration-tests/extensions-reload.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..4a1250fd00fe9180c37b32d4ca68aeeda845d883
--- /dev/null
+++ b/integration-tests/extensions-reload.test.ts
@@ -0,0 +1,151 @@
+/**
+ * @license
+ * Copyright 2025 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { expect, it, describe, beforeEach, afterEach } from 'vitest';
+import { TestRig } from './test-helper.js';
+import { TestMcpServer } from './test-mcp-server.js';
+import { writeFileSync } from 'node:fs';
+import { join } from 'node:path';
+import { safeJsonStringify } from '@google/gemini-cli-core/src/utils/safeJsonStringify.js';
+
+import stripAnsi from 'strip-ansi';
+
+describe('extension reloading', () => {
+ let rig: TestRig;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => await rig.cleanup());
+
+ // always fails
+ // TODO(#14527): Re-enable this once fixed
+ it.skip('installs a local extension, updates it, checks it was reloaded properly', async () => {
+ const serverA = new TestMcpServer();
+ const portA = await serverA.start({
+ hello: () => ({ content: [{ type: 'text', text: 'world' }] }),
+ });
+ const extension = {
+ name: 'test-extension',
+ version: '0.0.1',
+ mcpServers: {
+ 'test-server': {
+ httpUrl: `http://localhost:${portA}/mcp`,
+ },
+ },
+ };
+
+ rig.setup('extension reload test', {
+ settings: {
+ experimental: { extensionReloading: true },
+ },
+ });
+ const testServerPath = join(rig.testDir!, 'gemini-extension.json');
+ writeFileSync(testServerPath, safeJsonStringify(extension, 2));
+ // defensive cleanup from previous tests.
+ try {
+ await rig.runCommand(['extensions', 'uninstall', 'test-extension']);
+ } catch {
+ /* empty */
+ }
+
+ const result = await rig.runCommand(
+ ['--debug', 'extensions', 'install', `${rig.testDir!}`],
+ { stdin: 'y\n' },
+ );
+ expect(result).toContain('test-extension');
+
+ // Now create the update, but its not installed yet
+ const serverB = new TestMcpServer();
+ const portB = await serverB.start({
+ goodbye: () => ({ content: [{ type: 'text', text: 'world' }] }),
+ });
+ extension.version = '0.0.2';
+ extension.mcpServers['test-server'].httpUrl =
+ `http://localhost:${portB}/mcp`;
+ writeFileSync(testServerPath, safeJsonStringify(extension, 2));
+
+ // Start the CLI.
+ const run = await rig.runInteractive({ args: '--debug' });
+ await run.expectText('You have 1 extension with an update available');
+ // See the outdated extension
+ await run.sendText('/extensions list');
+ await run.type('\r');
+ await run.expectText('test-extension (v0.0.1) - active (update available)');
+ // Wait for the UI to settle and retry the command until we see the update
+ await new Promise((resolve) => setTimeout(resolve, 1000));
+
+ // Poll for the updated list
+ await rig.pollCommand(
+ async () => {
+ await run.sendText('/mcp list');
+ await run.type('\r');
+ },
+ () => {
+ const output = stripAnsi(run.output);
+ return (
+ output.includes(
+ 'test-server (from test-extension) - Ready (1 tool)',
+ ) && output.includes('- mcp_test-server_hello')
+ );
+ },
+ 30000, // 30s timeout
+ );
+
+ // Update the extension, expect the list to update, and mcp servers as well.
+ await run.sendKeys('\u0015/extensions update test-extension');
+ await run.expectText('/extensions update test-extension');
+ await run.type('\r');
+ await new Promise((resolve) => setTimeout(resolve, 500));
+ await run.type('\r');
+ await run.expectText(
+ ` * test-server (remote): http://localhost:${portB}/mcp`,
+ );
+ await run.type('\r'); // consent
+ await run.expectText(
+ 'Extension "test-extension" successfully updated: 0.0.1 β 0.0.2',
+ );
+
+ // Poll for the updated extension version
+ await rig.pollCommand(
+ async () => {
+ await run.sendText('/extensions list');
+ await run.type('\r');
+ },
+ () =>
+ stripAnsi(run.output).includes(
+ 'test-extension (v0.0.2) - active (updated)',
+ ),
+ 30000,
+ );
+
+ // Poll for the updated mcp tool
+ await rig.pollCommand(
+ async () => {
+ await run.sendText('/mcp list');
+ await run.type('\r');
+ },
+ () => {
+ const output = stripAnsi(run.output);
+ return (
+ output.includes(
+ 'test-server (from test-extension) - Ready (1 tool)',
+ ) && output.includes('- mcp_test-server_goodbye')
+ );
+ },
+ 30000,
+ );
+
+ await run.sendText('/quit');
+ await run.type('\r');
+
+ // Clean things up.
+ await serverA.stop();
+ await serverB.stop();
+ await rig.runCommand(['extensions', 'uninstall', 'test-extension']);
+ });
+});
diff --git a/integration-tests/file-system-interactive.test.ts b/integration-tests/file-system-interactive.test.ts
new file mode 100644
index 0000000000000000000000000000000000000000..8d90b8a6779cc33d65eaf733425409c357f0b127
--- /dev/null
+++ b/integration-tests/file-system-interactive.test.ts
@@ -0,0 +1,67 @@
+/**
+ * @license
+ * Copyright 2025 Google LLC
+ * SPDX-License-Identifier: Apache-2.0
+ */
+
+import { expect, describe, it, beforeEach, afterEach } from 'vitest';
+import { TestRig, skipFlaky } from './test-helper.js';
+
+describe.skipIf(skipFlaky)('Interactive file system', () => {
+ let rig: TestRig;
+
+ beforeEach(() => {
+ rig = new TestRig();
+ });
+
+ afterEach(async () => {
+ await rig.cleanup();
+ });
+
+ it('should perform a read-then-write sequence', async () => {
+ const fileName = 'version.txt';
+ await rig.setup('interactive-read-then-write', {
+ settings: {
+ security: {
+ auth: {
+ selectedType: 'gemini-api-key',
+ },
+ disableYoloMode: false,
+ },
+ },
+ });
+ rig.createFile(fileName, '1.0.0');
+
+ const run = await rig.runInteractive({
+ env: {
+ GEMINI_CLI_TRUST_WORKSPACE: 'true',
+ },
+ });
+
+ // Step 1: Read the file
+ const readPrompt = `Read the version from ${fileName} using the read_file tool`;
+ await run.type(readPrompt);
+ await run.type('\r');
+
+ const readCall = await rig.waitForToolCall('read_file', 30000);
+ expect(readCall, 'Expected to find a read_file tool call').toBe(true);
+
+ // Wait for the CLI to finish outputting the response and show the prompt again
+ await run.expectText('Type your message', 30000);
+
+ // Step 2: Write the file
+ const writePrompt = `now change the version to 1.0.1 in ${fileName} using the write_file tool`;
+ await run.type(writePrompt);
+ await run.type('\r');
+
+ // Check tool calls made with right args
+ await rig.expectToolCallSuccess(
+ ['write_file', 'replace'],
+ 30000,
+ (args) => args.includes('1.0.1') && args.includes(fileName),
+ );
+
+ // Wait for telemetry to flush and file system to sync, especially in sandboxed environments
+ await rig.waitForTelemetryReady();
+ }, 120000);
+});
diff --git a/integration-tests/flicker-detector.max-height.responses b/integration-tests/flicker-detector.max-height.responses
new file mode 100644
index 0000000000000000000000000000000000000000..b905b3d866c4cf13bd3fac7a54187b9ef81f2186
--- /dev/null
+++ b/integration-tests/flicker-detector.max-height.responses
@@ -0,0 +1,3 @@
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"{\n \"reasoning\": \"The user is asking for a simple piece of information ('a fun fact'). This is a direct, bounded request with low operational complexity and does not require strategic planning, extensive investigation, or debugging.\",\n \"model_choice\": \"flash\"\n}"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":1173,"candidatesTokenCount":59,"totalTokenCount":1344,"promptTokensDetails":[{"modality":"TEXT","tokenCount":1173}],"thoughtsTokenCount":112}}}
+{"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"thought":true,"text":"**Locating a fun fact**\n\nI'm now searching for a fun fact using the web search tool, focusing on finding something engaging and potentially surprising. The goal is to provide a brief, interesting piece of information.\n\n\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12226,"totalTokenCount":12255,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12226}],"thoughtsTokenCount":29}},{"candidates":[{"content":{"parts":[{"thoughtSignature":"CikB0e2Kb1vYSbIdmBfclWY7z4mOZgPxUGi3CtNXYYV9CSmG+SpVXZZkmQpZAdHtim9HVruyrUZZcHKDvIfn3j6/zLMgepC4Pqd79pG641PkPJnnCqEfVFRxmE2NX3Tj2lwRhtuIYT9Cc3CfvWGjbuuvwzynMCApxpIvxdXac/fXJYeRHTsKQQHR7Ypv6eOvWUFUTRGm1x29v8ZnGjtudG31H/Dgc65Y47c594ZJfX9RqJJil0I52Bxsm8UQ74rbARqwT7zYEbNO","functionCall":{"name":"google_web_search","args":{"query":"fun fact"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12226,"candidatesTokenCount":17,"totalTokenCount":12272,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12226}],"thoughtsTokenCount":29}}]}
+{"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Here's a fun fact: A day on Venus is longer than a year on Venus. It takes approximately 243 Earth days for Venus to rotate once on its axis, while its orbit around the Sun is about 225 Earth days."}],"role":"model"},"finishReason":"STOP","groundingMetadata":{"searchEntryPoint":{"renderedContent":"\n