SaylorTwift HF Staff commited on
Commit
c44600e
·
verified ·
1 Parent(s): 739fc44

Add files using upload-large-folder tool

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. docs/CONTRIBUTING.md +571 -0
  2. docs/behavioral-evals.md +185 -0
  3. docs/index.md +139 -0
  4. docs/integration-tests.md +293 -0
  5. docs/issue-and-pr-automation.md +203 -0
  6. docs/local-development.md +182 -0
  7. docs/npm.md +62 -0
  8. docs/redirects.json +21 -0
  9. docs/release-confidence.md +168 -0
  10. docs/releases.md +550 -0
  11. docs/sidebar.json +298 -0
  12. integration-tests/acp-env-auth.test.ts +163 -0
  13. integration-tests/acp-telemetry.test.ts +116 -0
  14. integration-tests/api-resilience.responses +1 -0
  15. integration-tests/api-resilience.test.ts +50 -0
  16. integration-tests/browser-agent-localhost.multistep.responses +9 -0
  17. integration-tests/browser-agent-localhost.navigate.responses +5 -0
  18. integration-tests/browser-agent.cleanup.responses +5 -0
  19. integration-tests/browser-agent.confirmation.responses +1 -0
  20. integration-tests/browser-policy.responses +5 -0
  21. integration-tests/browser-policy.test.ts +240 -0
  22. integration-tests/checkpointing.test.ts +155 -0
  23. integration-tests/concurrency-limit.responses +12 -0
  24. integration-tests/context-compress-interactive.compress-empty.responses +0 -0
  25. integration-tests/context-compress-interactive.compress-failure.responses +3 -0
  26. integration-tests/context-fidelity.test.ts +287 -0
  27. integration-tests/extensions-install.test.ts +62 -0
  28. integration-tests/extensions-reload.test.ts +151 -0
  29. integration-tests/file-system-interactive.test.ts +67 -0
  30. integration-tests/flicker-detector.max-height.responses +3 -0
  31. integration-tests/globalSetup.ts +155 -0
  32. integration-tests/google_web_search.test.ts +95 -0
  33. integration-tests/hooks-agent-flow-multistep.responses +2 -0
  34. integration-tests/hooks-agent-flow.test.ts +338 -0
  35. integration-tests/hooks-system.after-agent.responses +3 -0
  36. integration-tests/hooks-system.after-model.responses +1 -0
  37. integration-tests/hooks-system.after-tool-context.responses +2 -0
  38. integration-tests/hooks-system.allow-tool.responses +4 -0
  39. integration-tests/hooks-system.before-model.responses +1 -0
  40. integration-tests/hooks-system.before-tool-stop.responses +1 -0
  41. integration-tests/hooks-system.compress-auto.responses +1 -0
  42. integration-tests/hooks-system.input-modification.responses +2 -0
  43. integration-tests/hooks-system.input-validation.responses +2 -0
  44. integration-tests/hooks-system.multiple-events.responses +4 -0
  45. integration-tests/hooks-system.notification.responses +1 -0
  46. integration-tests/hooks-system.sequential-execution.responses +1 -0
  47. integration-tests/hooks-system.session-clear.responses +1 -0
  48. integration-tests/hooks-system.session-startup.responses +1 -0
  49. integration-tests/hooks-system.tail-tool-call.responses +2 -0
  50. integration-tests/hooks-system.telemetry.responses +2 -0
docs/CONTRIBUTING.md ADDED
@@ -0,0 +1,571 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # How to contribute
2
+
3
+ We would love to accept your patches and contributions to this project. This
4
+ document includes:
5
+
6
+ - **[Before you begin](#before-you-begin):** Essential steps to take before
7
+ becoming a Gemini CLI contributor.
8
+ - **[Code contribution process](#code-contribution-process):** How to contribute
9
+ code to Gemini CLI.
10
+ - **[Development setup and workflow](#development-setup-and-workflow):** How to
11
+ set up your development environment and workflow.
12
+ - **[Documentation contribution process](#documentation-contribution-process):**
13
+ How to contribute documentation to Gemini CLI.
14
+
15
+ We're looking forward to seeing your contributions!
16
+
17
+ ## Before you begin
18
+
19
+ ### Sign our Contributor License Agreement
20
+
21
+ Contributions to this project must be accompanied by a
22
+ [Contributor License Agreement](https://cla.developers.google.com/about) (CLA).
23
+ You (or your employer) retain the copyright to your contribution; this simply
24
+ gives us permission to use and redistribute your contributions as part of the
25
+ project.
26
+
27
+ If you or your current employer have already signed the Google CLA (even if it
28
+ was for a different project), you probably don't need to do it again.
29
+
30
+ Visit <https://cla.developers.google.com/> to see your current agreements or to
31
+ sign a new one.
32
+
33
+ ### Review our Community Guidelines
34
+
35
+ This project follows
36
+ [Google's Open Source Community Guidelines](https://opensource.google/conduct/).
37
+
38
+ ## Code contribution process
39
+
40
+ ### Get started
41
+
42
+ The process for contributing code is as follows:
43
+
44
+ 1. **Find an issue** that you want to work on. If an issue is tagged as
45
+ `🔒Maintainers only`, this means it is reserved for project maintainers. We
46
+ will not accept pull requests related to these issues. In the near future,
47
+ we will explicitly mark issues looking for contributions using the
48
+ `help-wanted` label. If you believe an issue is a good candidate for
49
+ community contribution, please leave a comment on the issue. A maintainer
50
+ will review it and apply the `help-wanted` label if appropriate. Only
51
+ maintainers should attempt to add the `help-wanted` label to an issue.
52
+ 2. **Fork the repository** and create a new branch.
53
+ 3. **Make your changes** in the `packages/` directory.
54
+ 4. **Ensure all checks pass** by running `npm run preflight`.
55
+ 5. **Open a pull request** with your changes.
56
+
57
+ ### Code reviews
58
+
59
+ All submissions, including submissions by project members, require review. We
60
+ use [GitHub pull requests](https://docs.github.com/articles/about-pull-requests)
61
+ for this purpose.
62
+
63
+ To assist with the review process, we provide an automated review tool that
64
+ helps detect common anti-patterns, testing issues, and other best practices that
65
+ are easy to miss.
66
+
67
+ #### Using the automated review tool
68
+
69
+ You can run the review tool in two ways:
70
+
71
+ 1. **Using the helper script (Recommended):** We provide a script that
72
+ automatically handles checking out the PR into a separate worktree,
73
+ installing dependencies, building the project, and launching the review
74
+ tool.
75
+
76
+ ```bash
77
+ ./scripts/review.sh <PR_NUMBER> [model]
78
+ ```
79
+
80
+ **Warning:** If you run `scripts/review.sh`, you must have first verified
81
+ that the code for the PR being reviewed is safe to run and does not contain
82
+ data exfiltration attacks.
83
+
84
+ **Authors are strongly encouraged to run this script on their own PRs**
85
+ immediately after creation. This allows you to catch and fix simple issues
86
+ locally before a maintainer performs a full review.
87
+
88
+ **Note on Models:** By default, the script uses the latest Pro model
89
+ (`gemini-3.1-pro-preview`). If you do not have enough Pro quota, you can run
90
+ it with the latest Flash model instead:
91
+ `./scripts/review.sh <PR_NUMBER> gemini-3-flash-preview`.
92
+
93
+ 2. **Manually from within Gemini CLI:** If you already have the PR checked out
94
+ and built, you can run the tool directly from the CLI prompt:
95
+
96
+ ```text
97
+ /review-frontend <PR_NUMBER>
98
+ ```
99
+
100
+ Replace `<PR_NUMBER>` with your pull request number. Reviewers should use this
101
+ tool to augment, not replace, their manual review process.
102
+
103
+ ### Self-assigning and unassigning issues
104
+
105
+ To assign an issue to yourself, simply add a comment with the text `/assign`. To
106
+ unassign yourself from an issue, add a comment with the text `/unassign`.
107
+
108
+ The comment must contain only that text and nothing else. These commands will
109
+ assign or unassign the issue as requested, provided the conditions are met
110
+ (e.g., an issue must be unassigned to be assigned).
111
+
112
+ Please note that you can have a maximum of 3 issues assigned to you at any given
113
+ time and that only
114
+ [issues labeled "help wanted"](https://github.com/google-gemini/gemini-cli/issues?q=is%3Aissue%20state%3Aopen%20label%3A%22help%20wanted%22)
115
+ may be self-assigned.
116
+
117
+ ### Pull request guidelines
118
+
119
+ To help us review and merge your PRs quickly, please follow these guidelines.
120
+ PRs that do not meet these standards may be closed.
121
+
122
+ #### 1. Link to an existing issue
123
+
124
+ All PRs should be linked to an existing issue in our tracker. This ensures that
125
+ every change has been discussed and is aligned with the project's goals before
126
+ any code is written.
127
+
128
+ - **For bug fixes:** The PR should be linked to the bug report issue.
129
+ - **For features:** The PR should be linked to the feature request or proposal
130
+ issue that has been approved by a maintainer.
131
+
132
+ If an issue for your change doesn't exist, we will automatically close your PR
133
+ along with a comment reminding you to associate the PR with an issue. The ideal
134
+ workflow starts with an issue that has been reviewed and approved by a
135
+ maintainer. Please **open the issue first** and wait for feedback before you
136
+ start coding.
137
+
138
+ #### 2. Keep it small and focused
139
+
140
+ We favor small, atomic PRs that address a single issue or add a single,
141
+ self-contained feature.
142
+
143
+ - **Do:** Create a PR that fixes one specific bug or adds one specific feature.
144
+ - **Don't:** Bundle multiple unrelated changes (e.g., a bug fix, a new feature,
145
+ and a refactor) into a single PR.
146
+
147
+ Large changes should be broken down into a series of smaller, logical PRs that
148
+ can be reviewed and merged independently.
149
+
150
+ #### 3. Use draft PRs for work in progress
151
+
152
+ If you'd like to get early feedback on your work, please use GitHub's **Draft
153
+ Pull Request** feature. This signals to the maintainers that the PR is not yet
154
+ ready for a formal review but is open for discussion and initial feedback.
155
+
156
+ #### 4. Ensure all checks pass
157
+
158
+ Before submitting your PR, ensure that all automated checks are passing by
159
+ running `npm run preflight`. This command runs all tests, linting, and other
160
+ style checks.
161
+
162
+ #### 5. Update documentation
163
+
164
+ If your PR introduces a user-facing change (e.g., a new command, a modified
165
+ flag, or a change in behavior), you must also update the relevant documentation
166
+ in the `/docs` directory.
167
+
168
+ See more about writing documentation:
169
+ [Documentation contribution process](#documentation-contribution-process).
170
+
171
+ #### 6. Write clear commit messages and a good PR description
172
+
173
+ Your PR should have a clear, descriptive title and a detailed description of the
174
+ changes. Follow the [Conventional Commits](https://www.conventionalcommits.org/)
175
+ standard for your commit messages.
176
+
177
+ - **Good PR title:** `feat(cli): Add --json flag to 'config get' command`
178
+ - **Bad PR title:** `Made some changes`
179
+
180
+ In the PR description, explain the "why" behind your changes and link to the
181
+ relevant issue (e.g., `Fixes #123`).
182
+
183
+ ### Forking
184
+
185
+ If you are forking the repository you will be able to run the Build, Test and
186
+ Integration test workflows. However in order to make the integration tests run
187
+ you'll need to add a
188
+ [GitHub Repository Secret](https://docs.github.com/en/actions/security-for-github-actions/security-guides/using-secrets-in-github-actions#creating-secrets-for-a-repository)
189
+ with a value of `GEMINI_API_KEY` and set that to a valid API key that you have
190
+ available. Your key and secret are private to your repo; no one without access
191
+ can see your key and you cannot see any secrets related to this repo.
192
+
193
+ Additionally you will need to click on the `Actions` tab and enable workflows
194
+ for your repository, you'll find it's the large blue button in the center of the
195
+ screen.
196
+
197
+ ### Development setup and workflow
198
+
199
+ This section guides contributors on how to build, modify, and understand the
200
+ development setup of this project.
201
+
202
+ ### Setting up the development environment
203
+
204
+ **Prerequisites:**
205
+
206
+ 1. **Node.js**:
207
+ - **Development:** Please use Node.js `~20.19.0`. This specific version is
208
+ required due to an upstream development dependency issue. You can use a
209
+ tool like [nvm](https://github.com/nvm-sh/nvm) to manage Node.js versions.
210
+ - **Production:** For running the CLI in a production environment, any
211
+ version of Node.js `>=20` is acceptable.
212
+ 2. **Git**
213
+
214
+ ### Build process
215
+
216
+ To clone the repository:
217
+
218
+ ```bash
219
+ git clone https://github.com/google-gemini/gemini-cli.git # Or your fork's URL
220
+ cd gemini-cli
221
+ ```
222
+
223
+ To install dependencies defined in `package.json` as well as root dependencies:
224
+
225
+ ```bash
226
+ npm install
227
+ ```
228
+
229
+ To build the entire project (all packages):
230
+
231
+ ```bash
232
+ npm run build
233
+ ```
234
+
235
+ This command typically compiles TypeScript to JavaScript, bundles assets, and
236
+ prepares the packages for execution. Refer to `scripts/build.js` and
237
+ `package.json` scripts for more details on what happens during the build.
238
+
239
+ ### Enabling sandboxing
240
+
241
+ [Sandboxing](#sandboxing) is highly recommended and requires, at a minimum,
242
+ setting `GEMINI_SANDBOX=true` in your `~/.env` and ensuring a sandboxing
243
+ provider (e.g. `macOS Seatbelt`, `docker`, or `podman`) is available. See
244
+ [Sandboxing](#sandboxing) for details.
245
+
246
+ To build both the `gemini` CLI utility and the sandbox container, run
247
+ `build:all` from the root directory:
248
+
249
+ ```bash
250
+ npm run build:all
251
+ ```
252
+
253
+ To skip building the sandbox container, you can use `npm run build` instead.
254
+
255
+ ### Running the CLI
256
+
257
+ To start the Gemini CLI from the source code (after building), run the following
258
+ command from the root directory:
259
+
260
+ ```bash
261
+ npm start
262
+ ```
263
+
264
+ If you'd like to run the source build outside of the gemini-cli folder, you can
265
+ utilize `npm link path/to/gemini-cli/packages/cli` (see:
266
+ [docs](https://docs.npmjs.com/cli/v9/commands/npm-link)) or
267
+ `alias gemini="node path/to/gemini-cli/packages/cli"` to run with `gemini`
268
+
269
+ ### Running tests
270
+
271
+ This project contains two types of tests: unit tests and integration tests.
272
+
273
+ #### Unit tests
274
+
275
+ To execute the unit test suite for the project:
276
+
277
+ ```bash
278
+ npm run test
279
+ ```
280
+
281
+ This will run tests located in the `packages/core` and `packages/cli`
282
+ directories. Ensure tests pass before submitting any changes. For a more
283
+ comprehensive check, it is recommended to run `npm run preflight`.
284
+
285
+ #### Integration tests
286
+
287
+ The integration tests are designed to validate the end-to-end functionality of
288
+ the Gemini CLI. They are not run as part of the default `npm run test` command.
289
+
290
+ To run the integration tests, use the following command:
291
+
292
+ ```bash
293
+ npm run test:e2e
294
+ ```
295
+
296
+ For more detailed information on the integration testing framework, please see
297
+ the
298
+ [Integration Tests documentation](https://geminicli.com/docs/integration-tests).
299
+
300
+ ### Linting and preflight checks
301
+
302
+ To ensure code quality and formatting consistency, run the preflight check:
303
+
304
+ ```bash
305
+ npm run preflight
306
+ ```
307
+
308
+ This command will run ESLint, Prettier, all tests, and other checks as defined
309
+ in the project's `package.json`.
310
+
311
+ _ProTip_
312
+
313
+ after cloning create a git precommit hook file to ensure your commits are always
314
+ clean.
315
+
316
+ ```bash
317
+ echo "
318
+ # Run npm build and check for errors
319
+ if ! npm run preflight; then
320
+ echo "npm build failed. Commit aborted."
321
+ exit 1
322
+ fi
323
+ " > .git/hooks/pre-commit && chmod +x .git/hooks/pre-commit
324
+ ```
325
+
326
+ #### Formatting
327
+
328
+ To separately format the code in this project, run the following command from
329
+ the root directory:
330
+
331
+ ```bash
332
+ npm run format
333
+ ```
334
+
335
+ This command uses Prettier to format the code according to the project's style
336
+ guidelines.
337
+
338
+ #### Linting
339
+
340
+ To separately lint the code in this project, run the following command from the
341
+ root directory:
342
+
343
+ ```bash
344
+ npm run lint
345
+ ```
346
+
347
+ ### Coding conventions
348
+
349
+ - Please adhere to the coding style, patterns, and conventions used throughout
350
+ the existing codebase.
351
+ - Consult
352
+ [GEMINI.md](https://github.com/google-gemini/gemini-cli/blob/main/GEMINI.md)
353
+ (typically found in the project root) for specific instructions related to
354
+ AI-assisted development, including conventions for React, comments, and Git
355
+ usage.
356
+ - **Imports:** Pay special attention to import paths. The project uses ESLint to
357
+ enforce restrictions on relative imports between packages.
358
+
359
+ ### Debugging
360
+
361
+ #### VS Code
362
+
363
+ 0. Run the CLI to interactively debug in VS Code with `F5`
364
+ 1. Start the CLI in debug mode from the root directory:
365
+ ```bash
366
+ npm run debug
367
+ ```
368
+ This command runs `node --inspect-brk dist/gemini.js` within the
369
+ `packages/cli` directory, pausing execution until a debugger attaches. You
370
+ can then open `chrome://inspect` in your Chrome browser to connect to the
371
+ debugger.
372
+ 2. In VS Code, use the "Attach" launch configuration (found in
373
+ `.vscode/launch.json`).
374
+
375
+ Alternatively, you can use the "Launch Program" configuration in VS Code if you
376
+ prefer to launch the currently open file directly, but 'F5' is generally
377
+ recommended.
378
+
379
+ To hit a breakpoint inside the sandbox container run:
380
+
381
+ ```bash
382
+ DEBUG=1 gemini
383
+ ```
384
+
385
+ **Note:** If you have `DEBUG=true` in a project's `.env` file, it won't affect
386
+ gemini-cli due to automatic exclusion. Use `.gemini/.env` files for gemini-cli
387
+ specific debug settings.
388
+
389
+ ### React DevTools
390
+
391
+ To debug the CLI's React-based UI, you can use React DevTools.
392
+
393
+ 1. **Start the Gemini CLI in development mode:**
394
+
395
+ ```bash
396
+ DEV=true npm start
397
+ ```
398
+
399
+ 2. **Install and run React DevTools version 6 (which matches the CLI's
400
+ `react-devtools-core`):**
401
+
402
+ You can either install it globally:
403
+
404
+ ```bash
405
+ npm install -g react-devtools@6
406
+ react-devtools
407
+ ```
408
+
409
+ Or run it directly using npx:
410
+
411
+ ```bash
412
+ npx react-devtools@6
413
+ ```
414
+
415
+ Your running CLI application should then connect to React DevTools.
416
+ ![](/docs/assets/connected_devtools.png)
417
+
418
+ ### Sandboxing
419
+
420
+ #### macOS Seatbelt
421
+
422
+ On macOS, `gemini` uses Seatbelt (`sandbox-exec`) under a `permissive-open`
423
+ profile (see `packages/cli/src/utils/sandbox-macos-permissive-open.sb`) that
424
+ denies operations by default, confining writes to the project folder while
425
+ allowing broad file reads and outbound network traffic ("open") by default. You
426
+ can switch to a `strict-open` profile (see
427
+ `packages/cli/src/utils/sandbox-macos-strict-open.sb`) that restricts both reads
428
+ and writes to the working directory while allowing outbound network traffic by
429
+ setting `SEATBELT_PROFILE=strict-open` in your environment or `.env` file.
430
+ Available built-in profiles are `permissive-{open,proxied}`,
431
+ `restrictive-{open,proxied}`, and `strict-{open,proxied}` (see below for proxied
432
+ networking). You can also switch to a custom profile
433
+ `SEATBELT_PROFILE=<profile>` if you also create a file
434
+ `.gemini/sandbox-macos-<profile>.sb` under your project settings directory
435
+ `.gemini`.
436
+
437
+ #### Container-based sandboxing (all platforms)
438
+
439
+ For stronger container-based sandboxing on macOS or other platforms, you can set
440
+ `GEMINI_SANDBOX=true|docker|podman|<command>` in your environment or `.env`
441
+ file. The specified command (or if `true` then either `docker` or `podman`) must
442
+ be installed on the host machine. Once enabled, `npm run build:all` will build a
443
+ minimal container ("sandbox") image and `npm start` will launch inside a fresh
444
+ instance of that container. The first build can take 20-30s (mostly due to
445
+ downloading of the base image) but after that both build and start overhead
446
+ should be minimal. Default builds (`npm run build`) will not rebuild the
447
+ sandbox.
448
+
449
+ Container-based sandboxing mounts the project directory (and system temp
450
+ directory) with read-write access and is started/stopped/removed automatically
451
+ as you start/stop Gemini CLI. Files created within the sandbox should be
452
+ automatically mapped to your user/group on host machine. You can easily specify
453
+ additional mounts, ports, or environment variables by setting
454
+ `SANDBOX_{MOUNTS,PORTS,ENV}` as needed. You can also fully customize the sandbox
455
+ for your projects by creating the files `.gemini/sandbox.Dockerfile` and/or
456
+ `.gemini/sandbox.bashrc` under your project settings directory (`.gemini`) and
457
+ running `gemini` with `BUILD_SANDBOX=1` to trigger building of your custom
458
+ sandbox.
459
+
460
+ #### Proxied networking
461
+
462
+ All sandboxing methods, including macOS Seatbelt using `*-proxied` profiles,
463
+ support restricting outbound network traffic through a custom proxy server that
464
+ can be specified as `GEMINI_SANDBOX_PROXY_COMMAND=<command>`, where `<command>`
465
+ must start a proxy server that listens on `:::8877` for relevant requests. See
466
+ `docs/examples/proxy-script.md` for a minimal proxy that only allows `HTTPS`
467
+ connections to `example.com:443` (e.g. `curl https://example.com`) and declines
468
+ all other requests. The proxy is started and stopped automatically alongside the
469
+ sandbox.
470
+
471
+ ### Manual publish
472
+
473
+ We publish an artifact for each commit to our internal registry. But if you need
474
+ to manually cut a local build, then run the following commands:
475
+
476
+ ```
477
+ npm run clean
478
+ npm install
479
+ npm run auth
480
+ npm run prerelease:dev
481
+ npm publish --workspaces
482
+ ```
483
+
484
+ ## Documentation contribution process
485
+
486
+ Our documentation must be kept up-to-date with our code contributions. We want
487
+ our documentation to be clear, concise, and helpful to our users. We value:
488
+
489
+ - **Clarity:** Use simple and direct language. Avoid jargon where possible.
490
+ - **Accuracy:** Ensure all information is correct and up-to-date.
491
+ - **Completeness:** Cover all aspects of a feature or topic.
492
+ - **Examples:** Provide practical examples to help users understand how to use
493
+ Gemini CLI.
494
+
495
+ ### Getting started
496
+
497
+ The process for contributing to the documentation is similar to contributing
498
+ code.
499
+
500
+ 1. **Fork the repository** and create a new branch.
501
+ 2. **Make your changes** in the `/docs` directory.
502
+ 3. **Preview your changes locally** in Markdown rendering.
503
+ 4. **Lint and format your changes.** Our preflight check includes linting and
504
+ formatting for documentation files.
505
+ ```bash
506
+ npm run preflight
507
+ ```
508
+ 5. **Open a pull request** with your changes.
509
+
510
+ ### Documentation structure
511
+
512
+ Our documentation is organized using
513
+ [sidebar.json](https://github.com/google-gemini/gemini-cli/blob/main/docs/sidebar.json)
514
+ as the table of contents. When adding new documentation:
515
+
516
+ 1. Create your markdown file **in the appropriate directory** under `/docs`.
517
+ 2. Add an entry to `sidebar.json` in the relevant section.
518
+ 3. Ensure all internal links use relative paths and point to existing files.
519
+
520
+ ### Style guide
521
+
522
+ We follow the
523
+ [Google Developer Documentation Style Guide](https://developers.google.com/style).
524
+ Please refer to it for guidance on writing style, tone, and formatting.
525
+
526
+ #### Key style points
527
+
528
+ - Use sentence case for headings.
529
+ - Write in second person ("you") when addressing the reader.
530
+ - Use present tense.
531
+ - Keep paragraphs short and focused.
532
+ - Use code blocks with appropriate language tags for syntax highlighting.
533
+ - Include practical examples whenever possible.
534
+
535
+ ### Linting and formatting
536
+
537
+ We use `prettier` to enforce a consistent style across our documentation. The
538
+ `npm run preflight` command will check for any linting issues.
539
+
540
+ You can also run the linter and formatter separately:
541
+
542
+ - `npm run lint` - Check for linting issues
543
+ - `npm run format` - Auto-format markdown files
544
+ - `npm run lint:fix` - Auto-fix linting issues where possible
545
+
546
+ Please make sure your contributions are free of linting errors before submitting
547
+ a pull request.
548
+
549
+ ### Before you submit
550
+
551
+ Before submitting your documentation pull request, please:
552
+
553
+ 1. Run `npm run preflight` to ensure all checks pass.
554
+ 2. Review your changes for clarity and accuracy.
555
+ 3. Check that all links work correctly.
556
+ 4. Ensure any code examples are tested and functional.
557
+ 5. Sign the
558
+ [Contributor License Agreement (CLA)](https://cla.developers.google.com/) if
559
+ you haven't already.
560
+
561
+ ### Need help?
562
+
563
+ If you have questions about contributing documentation:
564
+
565
+ - Check our [FAQ](https://geminicli.com/docs/resources/faq).
566
+ - Review existing documentation for examples.
567
+ - Open [an issue](https://github.com/google-gemini/gemini-cli/issues) to discuss
568
+ your proposed changes.
569
+ - Reach out to the maintainers.
570
+
571
+ We appreciate your contributions to making Gemini CLI documentation better!
docs/behavioral-evals.md ADDED
@@ -0,0 +1,185 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Behavioral Evaluations & EDK Guide
2
+
3
+ This guide introduces the **Eval Development Kit (EDK)** and details how to
4
+ write, validate, run, and report on **behavioral evaluations** in the Gemini CLI
5
+ codebase.
6
+
7
+ ---
8
+
9
+ ## Overview
10
+
11
+ Behavioral evaluations are automated tests designed to assert on the
12
+ **behavior** of the Gemini CLI agent (e.g., verifying which tools are called,
13
+ checking call ordering, or avoiding destructive commands) rather than checking
14
+ the final prose output.
15
+
16
+ Evaluating agent behavior is critical because:
17
+
18
+ 1. Model responses are non-deterministic, making exact prose matching highly
19
+ fragile.
20
+ 2. We must ensure the model utilizes the most efficient tools (e.g., batching
21
+ files via `read_many_files` instead of sequential `read_file` calls).
22
+ 3. We must enforce safety boundaries (e.g., preventing execution of raw shell
23
+ commands when safe alternatives exist).
24
+
25
+ All behavioral evaluations are stored under the `evals/` directory.
26
+
27
+ ---
28
+
29
+ ## EDK Developer Commands
30
+
31
+ The EDK provides CLI tools under `scripts/` to help contributors audit, check,
32
+ and monitor evals.
33
+
34
+ ### 1. `npm run eval:inventory`
35
+
36
+ Scans all eval files under `evals/`, statically parses them, and provides a
37
+ structured overview of what exists in the repository.
38
+
39
+ - **Usage:**
40
+ ```bash
41
+ npm run eval:inventory
42
+ ```
43
+ - **JSON Output:** For CI integration or inventory indexing, generate a
44
+ machine-readable JSON report:
45
+ ```bash
46
+ npm run eval:inventory -- --json
47
+ ```
48
+ - **Custom Root:** Run against another directory or repository:
49
+ ```bash
50
+ npm run eval:inventory -- --root /path/to/other/repo
51
+ ```
52
+
53
+ ---
54
+
55
+ ### 2. `npm run eval:validate`
56
+
57
+ A lint-like checker that validates eval source files against standard structural
58
+ guidelines and best practices.
59
+
60
+ - **Usage:**
61
+ ```bash
62
+ npm run eval:validate
63
+ ```
64
+ - **Custom Scopes:** Validate a specific file:
65
+ ```bash
66
+ npm run eval:validate -- evals/my-test.eval.ts
67
+ ```
68
+
69
+ #### Validation Rules & Severities
70
+
71
+ | Rule ID | Severity | Description |
72
+ | :------------------- | :---------- | :--------------------------------------------------------------------------------------------------------------------- |
73
+ | `file-naming` | **Error** | File must match `*.eval.ts` or `*.eval.tsx` naming conventions. |
74
+ | `valid-policy` | **Error** | Policy must be one of `ALWAYS_PASSES`, `USUALLY_PASSES`, or `USUALLY_FAILS`. |
75
+ | `suite-metadata` | **Error** | Both `suiteName` and `suiteType` must be present as static string literals. |
76
+ | `prompt-presence` | **Error** | Every eval case must have a non-empty `prompt` string. |
77
+ | `case-name-static` | **Error** | The case name must be a static string literal, not computed dynamically. |
78
+ | `invalid-tool-refs` | **Error** | All tools referenced in assertions must match known built-in or legacy tools. |
79
+ | `positive-assertion` | **Error** | Evaluation cases must assert on at least one tool call (e.g., check `waitForToolCall` has been invoked). |
80
+ | `workspace-setup` | **Error** | Workspace behaviors (like file-system edits/reads) must set up a `files` object. |
81
+ | `new-evals-policy` | **Warning** | New evals must not use `ALWAYS_PASSES` policy initially (they should be promoted after nightly data proves stability). |
82
+
83
+ Warnings (`new-evals-policy`) will be logged with `⚠` and will **not** cause
84
+ the CLI process to exit with status `1`. Errors (`✗`) will block CI builds and
85
+ return exit status `1`.
86
+
87
+ ---
88
+
89
+ ### 3. `npm run eval:report`
90
+
91
+ Aggregates local vitest `report.json` artifacts, maps them against inventory
92
+ policies, and summarizes the pass rates per model.
93
+
94
+ - **Usage:**
95
+ ```bash
96
+ npm run eval:report
97
+ ```
98
+ By default, it scans `evals/logs/` recursively for `report.json` files.
99
+ - **Specifying Directory:**
100
+ ```bash
101
+ npm run eval:report -- /path/to/logs
102
+ ```
103
+ - **JSON Output:**
104
+ ```bash
105
+ npm run eval:report -- --json
106
+ ```
107
+
108
+ ---
109
+
110
+ ## Contributor Workflow
111
+
112
+ When writing a new behavioral evaluation, adhere to this workflow to ensure
113
+ high-quality, non-flaky test runs.
114
+
115
+ ### Step-by-Step Guide
116
+
117
+ 1. **Identify the Target Behavior**: Determine which tool calls need
118
+ verification (e.g., `web_fetch` must be called).
119
+ 2. **Author the Eval File**: Create your file under `evals/<name>.eval.ts`
120
+ naming it properly.
121
+ 3. **Configure Workspace Files**: If the eval reads or edits files, define them
122
+ inside the `files` metadata field.
123
+ 4. **Assert Behavior, Not Prose**: Ensure the `assert` block checks tool
124
+ interactions using `rig.waitForToolCall` or similar. Do not check final
125
+ prose.
126
+ 5. **Run Locally**:
127
+ ```bash
128
+ RUN_EVALS=true npx vitest run evals/my-test.eval.ts
129
+ ```
130
+ 6. **Deflake**: Run your eval at least 3 times locally to verify it does not
131
+ fail due to model variance.
132
+ 7. **Run Validation**: Run `npm run eval:validate` to ensure no linting errors
133
+ are present.
134
+
135
+ ### Acceptance Criteria Checklist
136
+
137
+ - [ ] **Naming**: File ends with `.eval.ts` or `.eval.tsx`.
138
+ - [ ] **Policy**: New evals start as `USUALLY_PASSES`.
139
+ - [ ] **Metadata**: Static `suiteName` and `suiteType` (e.g. `'behavioral'`) are
140
+ specified.
141
+ - [ ] **Assertions**: Uses `rig.waitForToolCall` or asserts tool arguments
142
+ explicitly.
143
+ - [ ] **Clean workspace**: Does not write to files outside `rig.testDir`.
144
+
145
+ ### Common Anti-Patterns to Avoid
146
+
147
+ - **Restricting core tools**: Never override `settings.tools.core` to limit
148
+ tools. Evals must run against the default toolset.
149
+ - **Checking model prose**: Avoid `expect(result).toContain('something')` since
150
+ model wording is non-deterministic.
151
+ - **Integration-only testing**: Evals that only write files without checking
152
+ realistic model prompts are integration tests and belong under
153
+ `integration-tests/`.
154
+
155
+ ---
156
+
157
+ ## CI & Dashboard Integration
158
+
159
+ You can easily automate behavioral evaluations or compile dashboard data using
160
+ EDK's JSON reporters.
161
+
162
+ ### CI Validation Block
163
+
164
+ Add a step in your PR checks or GitHub workflows to automatically lint new evals
165
+ and block pull requests containing validation errors:
166
+
167
+ ```yaml
168
+ - name: Run Eval Validator
169
+ run: npm run eval:validate
170
+ ```
171
+
172
+ ### Publishing to a Dashboard
173
+
174
+ To record nightly performance metrics across multiple models:
175
+
176
+ 1. Configure your workflow to run evaluations with the JSON reporter:
177
+ ```bash
178
+ cross-env GEMINI_MODEL=gemini-2.5-pro npx vitest run --config evals/vitest.config.ts --reporter=json --outputFile="evals/logs/eval-logs-gemini-2.5-pro/report.json"
179
+ ```
180
+ 2. Aggregate all test runs using the reporting tool:
181
+ ```bash
182
+ npm run eval:report -- evals/logs --json > aggregated_report.json
183
+ ```
184
+ 3. Upload `aggregated_report.json` to your dashboard storage backend to
185
+ visualize pass rates over time.
docs/index.md ADDED
@@ -0,0 +1,139 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Gemini CLI documentation
2
+
3
+ Gemini CLI brings the power of Gemini models directly into your terminal. Use it
4
+ to understand code, automate tasks, and build workflows with your local project
5
+ context.
6
+
7
+ ## Install
8
+
9
+ ```bash
10
+ npm install -g @google/gemini-cli
11
+ ```
12
+
13
+ ## Get started
14
+
15
+ Jump in to Gemini CLI.
16
+
17
+ - **[Quickstart](./get-started/index.md):** Your first session with Gemini CLI.
18
+ - **[Installation](./get-started/installation.mdx):** How to install Gemini CLI
19
+ on your system.
20
+ - **[Authentication](./get-started/authentication.mdx):** Setup instructions for
21
+ personal and enterprise accounts.
22
+ - **[CLI cheatsheet](./cli/cli-reference.md):** A quick reference for common
23
+ commands and options.
24
+ - **[Gemini 3 on Gemini CLI](./get-started/gemini-3.md):** Learn about Gemini 3
25
+ support in Gemini CLI.
26
+
27
+ ## Use Gemini CLI
28
+
29
+ User-focused guides and tutorials for daily development workflows.
30
+
31
+ - **[File management](./cli/tutorials/file-management.md):** How to work with
32
+ local files and directories.
33
+ - **[Get started with Agent skills](./cli/tutorials/skills-getting-started.md):**
34
+ Getting started with specialized expertise.
35
+ - **[Manage context and memory](./cli/tutorials/memory-management.md):**
36
+ Managing persistent instructions and facts.
37
+ - **[Execute shell commands](./cli/tutorials/shell-commands.md):** Executing
38
+ system commands safely.
39
+ - **[Manage sessions and history](./cli/tutorials/session-management.md):**
40
+ Resuming, managing, and rewinding conversations.
41
+ - **[Plan tasks with todos](./cli/tutorials/task-planning.md):** Using todos for
42
+ complex workflows.
43
+ - **[Web search and fetch](./cli/tutorials/web-tools.md):** Searching and
44
+ fetching content from the web.
45
+ - **[Set up an MCP server](./cli/tutorials/mcp-setup.md):** Set up an MCP
46
+ server.
47
+ - **[Automate tasks](./cli/tutorials/automation.md):** Automate tasks.
48
+
49
+ ## Features
50
+
51
+ Technical documentation for each capability of Gemini CLI.
52
+
53
+ - **[Extensions](./extensions/index.md):** Extend Gemini CLI with new tools and
54
+ capabilities.
55
+ - **[Agent Skills](./cli/skills.md):** Use specialized agents for specific
56
+ tasks.
57
+ - **[Checkpointing](./cli/checkpointing.md):** Automatic session snapshots.
58
+ - **[Headless mode](./cli/headless.md):** Programmatic and scripting interface.
59
+ - **[Hooks](./hooks/index.md):** Customize Gemini CLI behavior with scripts.
60
+ - **[IDE integration](./ide-integration/index.md):** Integrate Gemini CLI with
61
+ your favorite IDE.
62
+ - **[MCP servers](./tools/mcp-server.md):** Connect to and use remote agents.
63
+ - **[Model routing](./cli/model-routing.md):** Automatic fallback resilience.
64
+ - **[Model selection](./cli/model.md):** Choose the best model for your needs.
65
+ - **[Plan mode 🔬](./cli/plan-mode.md):** Use a safe, read-only mode for
66
+ planning complex changes.
67
+ - **[Subagents 🔬](./core/subagents.md):** Using specialized agents for specific
68
+ tasks.
69
+ - **[Remote subagents 🔬](./core/remote-agents.md):** Connecting to and using
70
+ remote agents.
71
+ - **[Rewind](./cli/rewind.md):** Rewind and replay sessions.
72
+ - **[Sandboxing](./cli/sandbox.md):** Isolate tool execution.
73
+ - **[Settings](./cli/settings.md):** Full configuration reference.
74
+ - **[Telemetry](./cli/telemetry.md):** Usage and performance metric details.
75
+ - **[Token caching](./cli/token-caching.md):** Performance optimization.
76
+
77
+ ## Configuration
78
+
79
+ Settings and customization options for Gemini CLI.
80
+
81
+ - **[Custom commands](./cli/custom-commands.md):** Personalized shortcuts.
82
+ - **[Enterprise configuration](./cli/enterprise.md):** Professional environment
83
+ controls.
84
+ - **[Ignore files (.geminiignore)](./cli/gemini-ignore.md):** Exclusion pattern
85
+ reference.
86
+ - **[Model configuration](./cli/generation-settings.md):** Fine-tune generation
87
+ parameters like temperature and thinking budget.
88
+ - **[Project context (GEMINI.md)](./cli/gemini-md.md):** Technical hierarchy of
89
+ context files.
90
+ - **[System prompt override](./cli/system-prompt.md):** Instruction replacement
91
+ logic.
92
+ - **[Themes](./cli/themes.md):** UI personalization technical guide.
93
+ - **[Trusted folders](./cli/trusted-folders.md):** Security permission logic.
94
+
95
+ ## Reference
96
+
97
+ Deep technical documentation and API specifications.
98
+
99
+ - **[Command reference](./reference/commands.md):** Detailed slash command
100
+ guide.
101
+ - **[Configuration reference](./reference/configuration.md):** Settings and
102
+ environment variables.
103
+ - **[Keyboard shortcuts](./reference/keyboard-shortcuts.md):** Productivity
104
+ tips.
105
+ - **[Memory import processor](./reference/memport.md):** How Gemini CLI
106
+ processes memory from various sources.
107
+ - **[Policy engine](./reference/policy-engine.md):** Fine-grained execution
108
+ control.
109
+ - **[Tools reference](./reference/tools.md):** Information on how tools are
110
+ defined, registered, and used.
111
+
112
+ ## Resources
113
+
114
+ Support, release history, and legal information.
115
+
116
+ - **[FAQ](./resources/faq.md):** Answers to frequently asked questions.
117
+ - **[Quota and pricing](./resources/quota-and-pricing.md):** Limits and billing
118
+ details.
119
+ - **[Terms and privacy](./resources/tos-privacy.md):** Official notices and
120
+ terms.
121
+ - **[Troubleshooting](./resources/troubleshooting.md):** Common issues and
122
+ solutions.
123
+ - **[Uninstall](./resources/uninstall.md):** How to uninstall Gemini CLI.
124
+
125
+ ## Development
126
+
127
+ - **[Contribution guide](/docs/contributing):** How to contribute to Gemini CLI.
128
+ - **[Integration testing](./integration-tests.md):** Running integration tests.
129
+ - **[Issue and PR automation](./issue-and-pr-automation.md):** Automation for
130
+ issues and pull requests.
131
+ - **[Local development](./local-development.md):** Setting up a local
132
+ development environment.
133
+ - **[NPM package structure](./npm.md):** The structure of the NPM packages.
134
+
135
+ ## Releases
136
+
137
+ - **[Release notes](./changelogs/index.md):** Release notes for all versions.
138
+ - **[Stable release](./changelogs/latest.md):** The latest stable release.
139
+ - **[Preview release](./changelogs/preview.md):** The latest preview release.
docs/integration-tests.md ADDED
@@ -0,0 +1,293 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Integration tests
2
+
3
+ This document provides information about the integration testing framework used
4
+ in this project.
5
+
6
+ ## Overview
7
+
8
+ The integration tests are designed to validate the end-to-end functionality of
9
+ Gemini CLI. They execute the built binary in a controlled environment and verify
10
+ that it behaves as expected when interacting with the file system.
11
+
12
+ These tests are located in the `integration-tests` directory and are run using a
13
+ custom test runner.
14
+
15
+ ## Building the tests
16
+
17
+ Prior to running any integration tests, you need to create a release bundle that
18
+ you want to actually test:
19
+
20
+ ```bash
21
+ npm run bundle
22
+ ```
23
+
24
+ You must re-run this command after making any changes to the CLI source code,
25
+ but not after making changes to tests.
26
+
27
+ ## Running the tests
28
+
29
+ The integration tests are not run as part of the default `npm run test` command.
30
+ They must be run explicitly using the `npm run test:integration:all` script.
31
+
32
+ The integration tests can also be run using the following shortcut:
33
+
34
+ ```bash
35
+ npm run test:e2e
36
+ ```
37
+
38
+ ## Running a specific set of tests
39
+
40
+ To run a subset of test files, you can use
41
+ `npm run <integration test command> <file_name1> ....` where &lt;integration
42
+ test command&gt; is either `test:e2e` or `test:integration*` and `<file_name>`
43
+ is any of the `.test.js` files in the `integration-tests/` directory. For
44
+ example, the following command runs `list_directory.test.js` and
45
+ `write_file.test.js`:
46
+
47
+ ```bash
48
+ npm run test:e2e list_directory write_file
49
+ ```
50
+
51
+ ### Running a single test by name
52
+
53
+ To run a single test by its name, use the `--test-name-pattern` flag:
54
+
55
+ ```bash
56
+ npm run test:e2e -- --test-name-pattern "reads a file"
57
+ ```
58
+
59
+ ### Regenerating model responses
60
+
61
+ Some integration tests use faked out model responses, which may need to be
62
+ regenerated from time to time as the implementations change.
63
+
64
+ To regenerate these golden files, set the REGENERATE_MODEL_GOLDENS environment
65
+ variable to "true" when running the tests, for example:
66
+
67
+ **WARNING**: If running locally you should review these updated responses for
68
+ any information about yourself or your system that gemini may have included in
69
+ these responses.
70
+
71
+ ```bash
72
+ REGENERATE_MODEL_GOLDENS="true" npm run test:e2e
73
+ ```
74
+
75
+ **WARNING**: Make sure you run **await rig.cleanup()** at the end of your test,
76
+ else the golden files will not be updated.
77
+
78
+ ### Deflaking a test
79
+
80
+ Before adding a **new** integration test, you should test it at least 5 times
81
+ with the deflake script or workflow to make sure that it is not flaky.
82
+
83
+ ### Deflake script
84
+
85
+ ```bash
86
+ npm run deflake -- --runs=5 --command="npm run test:e2e -- -- --test-name-pattern '<your-new-test-name>'"
87
+ ```
88
+
89
+ #### Deflake workflow
90
+
91
+ ```bash
92
+ gh workflow run deflake.yml --ref <your-branch> -f test_name_pattern="<your-test-name-pattern>"
93
+ ```
94
+
95
+ ### Running all tests
96
+
97
+ To run the entire suite of integration tests, use the following command:
98
+
99
+ ```bash
100
+ npm run test:integration:all
101
+ ```
102
+
103
+ ### Sandbox matrix
104
+
105
+ The `all` command will run tests for `no sandboxing`, `docker` and `podman`.
106
+ Each individual type can be run using the following commands:
107
+
108
+ ```bash
109
+ npm run test:integration:sandbox:none
110
+ ```
111
+
112
+ ```bash
113
+ npm run test:integration:sandbox:docker
114
+ ```
115
+
116
+ ```bash
117
+ npm run test:integration:sandbox:podman
118
+ ```
119
+
120
+ ## Memory regression tests
121
+
122
+ Memory regression tests are designed to detect heap growth and leaks across key
123
+ CLI scenarios. They are located in the `memory-tests` directory.
124
+
125
+ These tests are distinct from standard integration tests because they measure
126
+ memory usage and compare it against committed baselines.
127
+
128
+ ### Running memory tests
129
+
130
+ Memory tests are not run as part of the default `npm run test` or
131
+ `npm run test:e2e` commands. They are run nightly in CI but can be run manually:
132
+
133
+ ```bash
134
+ npm run test:memory
135
+ ```
136
+
137
+ ### Updating baselines
138
+
139
+ If you intentionally change behavior that affects memory usage, you may need to
140
+ update the baselines. Set the `UPDATE_MEMORY_BASELINES` environment variable to
141
+ `true`:
142
+
143
+ ```bash
144
+ UPDATE_MEMORY_BASELINES=true npm run test:memory
145
+ ```
146
+
147
+ This will run the tests, take median snapshots, and overwrite
148
+ `memory-tests/baselines.json`. You should review the changes and commit the
149
+ updated baseline file.
150
+
151
+ ### How it works
152
+
153
+ The harness (`MemoryTestHarness` in `packages/test-utils`):
154
+
155
+ - Forces garbage collection multiple times to reduce noise.
156
+ - Takes median snapshots to filter spikes.
157
+ - Compares against baselines with a 10% tolerance.
158
+ - Can analyze sustained leaks across 3 snapshots using `analyzeSnapshots()`.
159
+
160
+ ## Performance regression tests
161
+
162
+ Performance regression tests are designed to detect wall-clock time, CPU usage,
163
+ and event loop delay regressions across key CLI scenarios. They are located in
164
+ the `perf-tests` directory.
165
+
166
+ These tests are distinct from standard integration tests because they measure
167
+ performance metrics and compare it against committed baselines.
168
+
169
+ ### Running performance tests
170
+
171
+ Performance tests are not run as part of the default `npm run test` or
172
+ `npm run test:e2e` commands. They are run nightly in CI but can be run manually:
173
+
174
+ ```bash
175
+ npm run test:perf
176
+ ```
177
+
178
+ ### Updating baselines
179
+
180
+ If you intentionally change behavior that affects performance, you may need to
181
+ update the baselines. Set the `UPDATE_PERF_BASELINES` environment variable to
182
+ `true`:
183
+
184
+ ```bash
185
+ UPDATE_PERF_BASELINES=true npm run test:perf
186
+ ```
187
+
188
+ This will run the tests multiple times (with warmup), apply IQR outlier
189
+ filtering, and overwrite `perf-tests/baselines.json`. You should review the
190
+ changes and commit the updated baseline file.
191
+
192
+ ### How it works
193
+
194
+ The harness (`PerfTestHarness` in `packages/test-utils`):
195
+
196
+ - Measures wall-clock time using `performance.now()`.
197
+ - Measures CPU usage using `process.cpuUsage()`.
198
+ - Monitors event loop delay using `perf_hooks.monitorEventLoopDelay()`.
199
+ - Applies IQR (Interquartile Range) filtering to remove outlier samples.
200
+ - Compares against baselines with a 15% tolerance.
201
+
202
+ ## Diagnostics
203
+
204
+ The integration test runner provides several options for diagnostics to help
205
+ track down test failures.
206
+
207
+ ### Keeping test output
208
+
209
+ You can preserve the temporary files created during a test run for inspection.
210
+ This is useful for debugging issues with file system operations.
211
+
212
+ To keep the test output set the `KEEP_OUTPUT` environment variable to `true`.
213
+
214
+ ```bash
215
+ KEEP_OUTPUT=true npm run test:integration:sandbox:none
216
+ ```
217
+
218
+ When output is kept, the test runner will print the path to the unique directory
219
+ for the test run.
220
+
221
+ ### Verbose output
222
+
223
+ For more detailed debugging, set the `VERBOSE` environment variable to `true`.
224
+
225
+ ```bash
226
+ VERBOSE=true npm run test:integration:sandbox:none
227
+ ```
228
+
229
+ When using `VERBOSE=true` and `KEEP_OUTPUT=true` in the same command, the output
230
+ is streamed to the console and also saved to a log file within the test's
231
+ temporary directory.
232
+
233
+ The verbose output is formatted to clearly identify the source of the logs:
234
+
235
+ ```
236
+ --- TEST: <log dir>:<test-name> ---
237
+ ... output from the gemini command ...
238
+ --- END TEST: <log dir>:<test-name> ---
239
+ ```
240
+
241
+ ## Linting and formatting
242
+
243
+ To ensure code quality and consistency, the integration test files are linted as
244
+ part of the main build process. You can also manually run the linter and
245
+ auto-fixer.
246
+
247
+ ### Running the linter
248
+
249
+ To check for linting errors, run the following command:
250
+
251
+ ```bash
252
+ npm run lint
253
+ ```
254
+
255
+ You can include the `:fix` flag in the command to automatically fix any fixable
256
+ linting errors:
257
+
258
+ ```bash
259
+ npm run lint:fix
260
+ ```
261
+
262
+ ## Directory structure
263
+
264
+ The integration tests create a unique directory for each test run inside the
265
+ `.integration-tests` directory. Within this directory, a subdirectory is created
266
+ for each test file, and within that, a subdirectory is created for each
267
+ individual test case.
268
+
269
+ This structure makes it easy to locate the artifacts for a specific test run,
270
+ file, or case.
271
+
272
+ ```
273
+ .integration-tests/
274
+ └── <run-id>/
275
+ └── <test-file-name>.test.js/
276
+ └── <test-case-name>/
277
+ ├── output.log
278
+ └── ...other test artifacts...
279
+ ```
280
+
281
+ ## Continuous integration
282
+
283
+ To ensure the integration tests are always run, a GitHub Actions workflow is
284
+ defined in `.github/workflows/chained_e2e.yml`. This workflow automatically runs
285
+ the integrations tests for pull requests against the `main` branch, or when a
286
+ pull request is added to a merge queue.
287
+
288
+ The workflow runs the tests in different sandboxing environments to ensure
289
+ Gemini CLI is tested across each:
290
+
291
+ - `sandbox:none`: Runs the tests without any sandboxing.
292
+ - `sandbox:docker`: Runs the tests in a Docker container.
293
+ - `sandbox:podman`: Runs the tests in a Podman container.
docs/issue-and-pr-automation.md ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Automation and triage processes
2
+
3
+ This document provides a detailed overview of the automated processes we use to
4
+ manage and triage issues and pull requests. Our goal is to provide prompt
5
+ feedback and ensure that contributions are reviewed and integrated efficiently.
6
+ Understanding this automation will help you as a contributor know what to expect
7
+ and how to best interact with our repository bots.
8
+
9
+ ## Guiding principle: Issues and pull requests
10
+
11
+ First and foremost, almost every Pull Request (PR) should be linked to a
12
+ corresponding Issue. The issue describes the "what" and the "why" (the bug or
13
+ feature), while the PR is the "how" (the implementation). This separation helps
14
+ us track work, prioritize features, and maintain clear historical context. Our
15
+ automation is built around this principle.
16
+
17
+ <!-- prettier-ignore -->
18
+ > [!NOTE]
19
+ > Issues tagged as "🔒Maintainers only" are reserved for project
20
+ > maintainers. We will not accept pull requests related to these issues.
21
+
22
+ ---
23
+
24
+ ## Detailed automation workflows
25
+
26
+ Here is a breakdown of the specific automation workflows that run in our
27
+ repository.
28
+
29
+ ### 1. When you open an issue: `Automated Issue Triage`
30
+
31
+ This is the first bot you will interact with when you create an issue. Its job
32
+ is to perform an initial analysis and apply the correct labels.
33
+
34
+ - **Workflow File**: `.github/workflows/gemini-automated-issue-triage.yml`
35
+ - **When it runs**: Immediately after an issue is created or reopened.
36
+ - **What it does**:
37
+ - It uses a Gemini model to analyze the issue's title and body against a
38
+ detailed set of guidelines.
39
+ - **Applies one `area/*` label**: Categorizes the issue into a functional area
40
+ of the project (for example, `area/ux`, `area/models`, `area/platform`).
41
+ - **Applies one `kind/*` label**: Identifies the type of issue (for example,
42
+ `kind/bug`, `kind/enhancement`, `kind/question`).
43
+ - **Applies one `priority/*` label**: Assigns a priority from P0 (critical) to
44
+ P3 (low) based on the described impact.
45
+ - **May apply `status/need-information`**: If the issue lacks critical details
46
+ (like logs or reproduction steps), it will be flagged for more information.
47
+ - **May apply `status/need-retesting`**: If the issue references a CLI version
48
+ that is more than six versions old, it will be flagged for retesting on a
49
+ current version.
50
+ - **What you should do**:
51
+ - Fill out the issue template as completely as possible. The more detail you
52
+ provide, the more accurate the triage will be.
53
+ - If the `status/need-information` label is added, provide the requested
54
+ details in a comment.
55
+
56
+ ### 2. When you open a pull request: `Continuous Integration (CI)`
57
+
58
+ This workflow ensures that all changes meet our quality standards before they
59
+ can be merged.
60
+
61
+ - **Workflow File**: `.github/workflows/ci.yml`
62
+ - **When it runs**: On every push to a pull request.
63
+ - **What it does**:
64
+ - **Lint**: Checks that your code adheres to our project's formatting and
65
+ style rules.
66
+ - **Test**: Runs our full suite of automated tests across macOS, Windows, and
67
+ Linux, and on multiple Node.js versions. This is the most time-consuming
68
+ part of the CI process.
69
+ - **Post Coverage Comment**: After all tests have successfully passed, a bot
70
+ will post a comment on your PR. This comment provides a summary of how well
71
+ your changes are covered by tests.
72
+ - **What you should do**:
73
+ - Ensure all CI checks pass. A green checkmark ✅ will appear next to your
74
+ commit when everything is successful.
75
+ - If a check fails (a red "X" ❌), click the "Details" link next to the failed
76
+ check to view the logs, identify the problem, and push a fix.
77
+
78
+ ### 3. Ongoing triage for pull requests: `PR Auditing and Label Sync`
79
+
80
+ This workflow runs periodically to ensure all open PRs are correctly linked to
81
+ issues and have consistent labels.
82
+
83
+ - **Workflow File**: `.github/workflows/gemini-scheduled-pr-triage.yml`
84
+ - **When it runs**: Every 15 minutes on all open pull requests.
85
+ - **What it does**:
86
+ - **Checks for a linked issue**: The bot scans your PR description for a
87
+ keyword that links it to an issue (for example, `Fixes #123`,
88
+ `Closes #456`).
89
+ - **Adds `status/need-issue`**: If no linked issue is found, the bot will add
90
+ the `status/need-issue` label to your PR. This is a clear signal that an
91
+ issue needs to be created and linked.
92
+ - **Synchronizes labels**: If an issue _is_ linked, the bot ensures the PR's
93
+ labels perfectly match the issue's labels. It will add any missing labels
94
+ and remove any that don't belong, and it will remove the `status/need-issue`
95
+ label if it was present.
96
+ - **What you should do**:
97
+ - **Always link your PR to an issue.** This is the most important step. Add a
98
+ line like `Resolves #<issue-number>` to your PR description.
99
+ - This will ensure your PR is correctly categorized and moves through the
100
+ review process smoothly.
101
+
102
+ ### 4. Ongoing triage for issues: `Scheduled Issue Triage`
103
+
104
+ This is a fallback workflow to ensure that no issue gets missed by the triage
105
+ process.
106
+
107
+ - **Workflow File**: `.github/workflows/gemini-scheduled-issue-triage.yml`
108
+ - **When it runs**: Every hour on all open issues.
109
+ - **What it does**:
110
+ - It actively seeks out issues that either have no labels at all or still have
111
+ the `status/need-triage` label.
112
+ - It then triggers the same powerful Gemini-based analysis as the initial
113
+ triage bot to apply the correct labels.
114
+ - **What you should do**:
115
+ - You typically don't need to do anything. This workflow is a safety net to
116
+ ensure every issue is eventually categorized, even if the initial triage
117
+ fails.
118
+
119
+ ### 5. Automatic unassignment of inactive contributors: `Unassign Inactive Issue Assignees`
120
+
121
+ To keep the list of open `help wanted` issues accessible to all contributors,
122
+ this workflow automatically removes **external contributors** who have not
123
+ opened a linked pull request within **7 days** of being assigned. Maintainers,
124
+ org members, and repo collaborators with write access or above are always exempt
125
+ and will never be auto-unassigned.
126
+
127
+ - **Workflow File**: `.github/workflows/unassign-inactive-assignees.yml`
128
+ - **When it runs**: Every day at 09:00 UTC, and can be triggered manually with
129
+ an optional `dry_run` mode.
130
+ - **What it does**:
131
+ 1. Finds every open issue labeled `help wanted` that has at least one
132
+ assignee.
133
+ 2. Identifies privileged users (team members, repo collaborators with write+
134
+ access, maintainers) and skips them entirely.
135
+ 3. For each remaining (external) assignee it reads the issue's timeline to
136
+ determine:
137
+ - The exact date they were assigned (using `assigned` timeline events).
138
+ - Whether they have opened a PR that is already linked/cross-referenced to
139
+ the issue.
140
+ 4. Each cross-referenced PR is fetched to verify it is **ready for review**:
141
+ open and non-draft, or already merged. Draft PRs do not count.
142
+ 5. If an assignee has been assigned for **more than 7 days** and no qualifying
143
+ PR is found, they are automatically unassigned and a comment is posted
144
+ explaining the reason and how to re-claim the issue.
145
+ 6. Assignees who have a non-draft, open or merged PR linked to the issue are
146
+ **never** unassigned by this workflow.
147
+ - **What you should do**:
148
+ - **Open a real PR, not a draft**: Within 7 days of being assigned, open a PR
149
+ that is ready for review and include `Fixes #<issue-number>` in the
150
+ description. Draft PRs do not satisfy the requirement and will not prevent
151
+ auto-unassignment.
152
+ - **Re-assign if unassigned by mistake**: Comment `/assign` on the issue to
153
+ assign yourself again.
154
+ - **Unassign yourself** if you can no longer work on the issue by commenting
155
+ `/unassign`, so other contributors can pick it up right away.
156
+
157
+ ### 6. Automatically label PRs by size: `PR Size Labeler`
158
+
159
+ To help maintainers estimate review effort and keep the PR history clean, this
160
+ workflow automatically tags every pull request with a size label representing
161
+ the total volume of line changes.
162
+
163
+ - **Workflow File**: `.github/workflows/pr-size-labeler.yml`
164
+ - **When it runs**: Immediately after a pull request is created, synchronized
165
+ (new commits pushed), or reopened. It can also be triggered manually via
166
+ `workflow_dispatch` with a PR number.
167
+ - **What it does**:
168
+ - **Calculates total changes**: Summarizes additions and deletions across all
169
+ changed files in a single consolidated API request.
170
+ - **Applies standard size labels**:
171
+ - `size/XS`: < 10 lines changed
172
+ - `size/S`: 10-49 lines changed
173
+ - `size/M`: 50-249 lines changed
174
+ - `size/L`: 250-999 lines changed
175
+ - `size/XL`: >= 1000 lines changed
176
+ - **Updates size tag atomically**: Adds the new correct size label and removes
177
+ any obsolete size labels in one atomic step.
178
+ - **Updates/Posts PR size info comment**: Instead of spamming a new comment on
179
+ every commit push, it updates the existing size labeler status comment
180
+ inline to keep the PR conversation timeline perfectly neat and clean.
181
+ - **What you should do**:
182
+ - You do not need to take any actions. The workflow runs automatically and
183
+ updates the label and comment seamlessly as you push new updates.
184
+
185
+ ### 7. Release automation
186
+
187
+ This workflow handles the process of packaging and publishing new versions of
188
+ Gemini CLI.
189
+
190
+ - **Workflow File**: `.github/workflows/release-manual.yml`
191
+ - **When it runs**: On a daily schedule for "nightly" releases, and manually for
192
+ official patch/minor releases.
193
+ - **What it does**:
194
+ - Automatically builds the project, bumps the version numbers, and publishes
195
+ the packages to npm.
196
+ - Creates a corresponding release on GitHub with generated release notes.
197
+ - **What you should do**:
198
+ - As a contributor, you don't need to do anything for this process. You can be
199
+ confident that once your PR is merged into the `main` branch, your changes
200
+ will be included in the very next nightly release.
201
+
202
+ We hope this detailed overview is helpful. If you have any questions about our
203
+ automation or processes, don't hesitate to ask!
docs/local-development.md ADDED
@@ -0,0 +1,182 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Local development guide
2
+
3
+ This guide provides instructions for setting up and using local development
4
+ features for Gemini CLI.
5
+
6
+ ## Tracing
7
+
8
+ Gemini CLI uses OpenTelemetry (OTel) to record traces that help you debug agent
9
+ behavior. Traces instrument key events like model calls, tool scheduler
10
+ operations, and tool calls.
11
+
12
+ Traces provide deep visibility into agent behavior and help you debug complex
13
+ issues. They are captured automatically when you enable telemetry.
14
+
15
+ ### View traces
16
+
17
+ You can view traces using Genkit Developer UI, Jaeger, or Google Cloud.
18
+
19
+ #### Use Genkit
20
+
21
+ Genkit provides a web-based UI for viewing traces and other telemetry data.
22
+
23
+ 1. **Start the Genkit telemetry server:**
24
+
25
+ Run the following command to start the Genkit server:
26
+
27
+ ```bash
28
+ npm run telemetry -- --target=genkit
29
+ ```
30
+
31
+ The script will output the URL for the Genkit Developer UI. For example:
32
+ `Genkit Developer UI: http://localhost:4000`
33
+
34
+ 2. **Run Gemini CLI:**
35
+
36
+ In a separate terminal, run your Gemini CLI command:
37
+
38
+ ```bash
39
+ gemini
40
+ ```
41
+
42
+ 3. **View the traces:**
43
+
44
+ Open the Genkit Developer UI URL in your browser and navigate to the
45
+ **Traces** tab to view the traces.
46
+
47
+ #### Use Jaeger
48
+
49
+ You can view traces in the Jaeger UI for local development.
50
+
51
+ 1. **Start the telemetry collector:**
52
+
53
+ Run the following command in your terminal to download and start Jaeger and
54
+ an OTel collector:
55
+
56
+ ```bash
57
+ npm run telemetry -- --target=local
58
+ ```
59
+
60
+ This command configures your workspace for local telemetry and provides a
61
+ link to the Jaeger UI (usually `http://localhost:16686`).
62
+
63
+ - **Collector logs:** `~/.gemini/tmp/<projectHash>/otel/collector.log`
64
+
65
+ 2. **Run Gemini CLI:**
66
+
67
+ In a separate terminal, run your Gemini CLI command:
68
+
69
+ ```bash
70
+ gemini
71
+ ```
72
+
73
+ 3. **View the traces:**
74
+
75
+ After running your command, open the Jaeger UI link in your browser to view
76
+ the traces.
77
+
78
+ #### Use Google Cloud
79
+
80
+ You can use an OpenTelemetry collector to forward telemetry data to Google Cloud
81
+ Trace for custom processing or routing.
82
+
83
+ <!-- prettier-ignore -->
84
+ > [!WARNING]
85
+ > Ensure you complete the
86
+ > [Google Cloud telemetry prerequisites](./cli/telemetry.md#prerequisites)
87
+ > (Project ID, authentication, IAM roles, and APIs) before using this method.
88
+
89
+ 1. **Configure `.gemini/settings.json`:**
90
+
91
+ ```json
92
+ {
93
+ "telemetry": {
94
+ "enabled": true,
95
+ "target": "gcp",
96
+ "useCollector": true
97
+ }
98
+ }
99
+ ```
100
+
101
+ 2. **Start the telemetry collector:**
102
+
103
+ Run the following command to start a local OTel collector that forwards to
104
+ Google Cloud:
105
+
106
+ ```bash
107
+ npm run telemetry -- --target=gcp
108
+ ```
109
+
110
+ The script outputs links to view traces, metrics, and logs in the Google
111
+ Cloud Console.
112
+
113
+ - **Collector logs:** `~/.gemini/tmp/<projectHash>/otel/collector-gcp.log`
114
+
115
+ 3. **Run Gemini CLI:**
116
+
117
+ In a separate terminal, run your Gemini CLI command:
118
+
119
+ ```bash
120
+ gemini
121
+ ```
122
+
123
+ 4. **View logs, metrics, and traces:**
124
+
125
+ After sending prompts, view your data in the Google Cloud Console. See the
126
+ [telemetry documentation](./cli/telemetry.md#view-google-cloud-telemetry)
127
+ for links to Logs, Metrics, and Trace explorers.
128
+
129
+ For more detailed information on telemetry, see the
130
+ [telemetry documentation](./cli/telemetry.md).
131
+
132
+ ### Instrument code with traces
133
+
134
+ You can add traces to your own code for more detailed instrumentation.
135
+
136
+ Adding traces helps you debug and understand the flow of execution. Use the
137
+ `runInDevTraceSpan` function to wrap any section of code in a trace span.
138
+
139
+ Here is a basic example:
140
+
141
+ ```typescript
142
+ import { runInDevTraceSpan } from '@google/gemini-cli-core';
143
+ import { GeminiCliOperation } from '@google/gemini-cli-core/lib/telemetry/constants.js';
144
+
145
+ await runInDevTraceSpan(
146
+ {
147
+ operation: GeminiCliOperation.ToolCall,
148
+ attributes: {
149
+ [GEN_AI_AGENT_NAME]: 'gemini-cli',
150
+ },
151
+ },
152
+ async ({ metadata }) => {
153
+ // metadata allows you to record the input and output of the
154
+ // operation as well as other attributes.
155
+ metadata.input = { key: 'value' };
156
+ // Set custom attributes.
157
+ metadata.attributes['custom.attribute'] = 'custom.value';
158
+
159
+ // Your code to be traced goes here.
160
+ try {
161
+ const output = await somethingRisky();
162
+ metadata.output = output;
163
+ return output;
164
+ } catch (e) {
165
+ metadata.error = e;
166
+ throw e;
167
+ }
168
+ },
169
+ );
170
+ ```
171
+
172
+ In this example:
173
+
174
+ - `operation`: The operation type of the span, represented by the
175
+ `GeminiCliOperation` enum.
176
+ - `metadata.input`: (Optional) An object containing the input data for the
177
+ traced operation.
178
+ - `metadata.output`: (Optional) An object containing the output data from the
179
+ traced operation.
180
+ - `metadata.attributes`: (Optional) A record of custom attributes to add to the
181
+ span.
182
+ - `metadata.error`: (Optional) An error object to record if the operation fails.
docs/npm.md ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Package overview
2
+
3
+ This monorepo contains two main packages: `@google/gemini-cli` and
4
+ `@google/gemini-cli-core`.
5
+
6
+ ## `@google/gemini-cli`
7
+
8
+ This is the main package for Gemini CLI. It is responsible for the user
9
+ interface, command parsing, and all other user-facing functionality.
10
+
11
+ When this package is published, it is bundled into a single executable file.
12
+ This bundle includes all of the package's dependencies, including
13
+ `@google/gemini-cli-core`. This means that whether a user installs the package
14
+ with `npm install -g @google/gemini-cli` or runs it directly with
15
+ `npx @google/gemini-cli`, they are using this single, self-contained executable.
16
+
17
+ ## `@google/gemini-cli-core`
18
+
19
+ This package contains the core logic for interacting with the Gemini API. It is
20
+ responsible for making API requests, handling authentication, and managing the
21
+ local cache.
22
+
23
+ This package is not bundled. When it is published, it is published as a standard
24
+ Node.js package with its own dependencies. This allows it to be used as a
25
+ standalone package in other projects, if needed. All transpiled js code in the
26
+ `dist` folder is included in the package.
27
+
28
+ ## NPM workspaces
29
+
30
+ This project uses
31
+ [NPM Workspaces](https://docs.npmjs.com/cli/v10/using-npm/workspaces) to manage
32
+ the packages within this monorepo. This simplifies development by allowing us to
33
+ manage dependencies and run scripts across multiple packages from the root of
34
+ the project.
35
+
36
+ ### How it works
37
+
38
+ The root `package.json` file defines the workspaces for this project:
39
+
40
+ ```json
41
+ {
42
+ "workspaces": ["packages/*"]
43
+ }
44
+ ```
45
+
46
+ This tells NPM that any folder inside the `packages` directory is a separate
47
+ package that should be managed as part of the workspace.
48
+
49
+ ### Benefits of workspaces
50
+
51
+ - **Simplified dependency management**: Running `npm install` from the root of
52
+ the project will install all dependencies for all packages in the workspace
53
+ and link them together. This means you don't need to run `npm install` in each
54
+ package's directory.
55
+ - **Automatic linking**: Packages within the workspace can depend on each other.
56
+ When you run `npm install`, NPM will automatically create symlinks between the
57
+ packages. This means that when you make changes to one package, the changes
58
+ are immediately available to other packages that depend on it.
59
+ - **Simplified script execution**: You can run scripts in any package from the
60
+ root of the project using the `--workspace` flag. For example, to run the
61
+ `build` script in the `cli` package, you can run
62
+ `npm run build --workspace @google/gemini-cli`.
docs/redirects.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "/docs/architecture": "/docs/cli/index",
3
+ "/docs/cli/commands": "/docs/reference/commands",
4
+ "/docs/cli": "/docs",
5
+ "/docs/cli/index": "/docs",
6
+ "/docs/cli/keyboard-shortcuts": "/docs/reference/keyboard-shortcuts",
7
+ "/docs/cli/uninstall": "/docs/resources/uninstall",
8
+ "/docs/core/concepts": "/docs",
9
+ "/docs/core/memport": "/docs/reference/memport",
10
+ "/docs/core/policy-engine": "/docs/reference/policy-engine",
11
+ "/docs/core/tools-api": "/docs/reference/tools",
12
+ "/docs/reference/tools-api": "/docs/reference/tools",
13
+ "/docs/faq": "/docs/resources/faq",
14
+ "/docs/get-started/configuration": "/docs/reference/configuration",
15
+ "/docs/get-started/configuration-v1": "/docs/reference/configuration",
16
+ "/docs/get-started/examples": "/docs/get-started/index",
17
+ "/docs/index": "/docs",
18
+ "/docs/quota-and-pricing": "/docs/resources/quota-and-pricing",
19
+ "/docs/tos-privacy": "/docs/resources/tos-privacy",
20
+ "/docs/troubleshooting": "/docs/resources/troubleshooting"
21
+ }
docs/release-confidence.md ADDED
@@ -0,0 +1,168 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Release confidence strategy
2
+
3
+ This document outlines the strategy for gaining confidence in every release of
4
+ Gemini CLI. It serves as a checklist and quality gate for release manager to
5
+ ensure we are shipping a high-quality product.
6
+
7
+ ## The goal
8
+
9
+ To answer the question, "Is this release _truly_ ready for our users?" with a
10
+ high degree of confidence, based on a holistic evaluation of automated signals,
11
+ manual verification, and data.
12
+
13
+ ## Level 1: Automated gates (must pass)
14
+
15
+ These are the baseline requirements. If any of these fail, the release is a
16
+ no-go.
17
+
18
+ ### 1. CI/CD health
19
+
20
+ All workflows in `.github/workflows/ci.yml` must pass on the `main` branch (for
21
+ nightly) or the release branch (for preview/stable).
22
+
23
+ - **Platforms:** Tests must pass on **Linux and macOS**.
24
+
25
+ - **Checks:**
26
+ - **Linting:** No linting errors (ESLint, Prettier, etc.).
27
+ - **Typechecking:** No TypeScript errors.
28
+ - **Unit Tests:** All unit tests in `packages/core` and `packages/cli` must
29
+ pass.
30
+ - **Build:** The project must build and bundle successfully.
31
+
32
+ ### 2. End-to-end (E2E) tests
33
+
34
+ All workflows in `.github/workflows/chained_e2e.yml` must pass.
35
+
36
+ - **Platforms:** **Linux, macOS and Windows**.
37
+ - **Sandboxing:** Tests must pass with both `sandbox:none` and `sandbox:docker`
38
+ on Linux.
39
+
40
+ ### 3. Post-deployment smoke tests
41
+
42
+ After a release is published to npm, the `smoke-test.yml` workflow runs. This
43
+ must pass to confirm the package is installable and the binary is executable.
44
+
45
+ - **Command:** `npx -y @google/gemini-cli@<tag> --version` must return the
46
+ correct version without error.
47
+ - **Platform:** Currently runs on `ubuntu-latest`.
48
+
49
+ ## Level 2: Manual verification and dogfooding
50
+
51
+ Automated tests cannot catch everything, especially UX issues.
52
+
53
+ ### 1. Dogfooding via `preview` tag
54
+
55
+ The weekly release cadence promotes code from `main` -> `nightly` -> `preview`
56
+ -> `stable`.
57
+
58
+ - **Requirement:** The `preview` release must be used by maintainers for at
59
+ least **one week** before being promoted to `stable`.
60
+ - **Action:** Maintainers should install the preview version locally:
61
+ ```bash
62
+ npm install -g @google/gemini-cli@preview
63
+ ```
64
+ - **Goal:** To catch regressions and UX issues in day-to-day usage before they
65
+ reach the broad user base.
66
+
67
+ ### 2. Critical user journey (CUJ) checklist
68
+
69
+ Before promoting a `preview` release to `stable`, a release manager must
70
+ manually run through this checklist.
71
+
72
+ - **Setup:**
73
+
74
+ - [ ] Uninstall any existing global version:
75
+ `npm uninstall -g @google/gemini-cli`
76
+ - [ ] Clear npx cache (optional but recommended): `npm cache clean --force`
77
+ - [ ] Install the preview version: `npm install -g @google/gemini-cli@preview`
78
+ - [ ] Verify version: `gemini --version`
79
+
80
+ - **Authentication:**
81
+
82
+ - [ ] In interactive mode run `/auth` and verify all sign in flows work:
83
+ - [ ] Sign in with Google
84
+ - [ ] API Key
85
+ - [ ] Vertex AI
86
+
87
+ - **Basic prompting:**
88
+
89
+ - [ ] Run `gemini "Tell me a joke"` and verify a sensible response.
90
+ - [ ] Run in interactive mode: `gemini`. Ask a follow-up question to test
91
+ context.
92
+
93
+ - **Piped input:**
94
+
95
+ - [ ] Run `echo "Summarize this" | gemini` and verify it processes stdin.
96
+
97
+ - **Context management:**
98
+
99
+ - [ ] In interactive mode, use `@file` to add a local file to context. Ask a
100
+ question about it.
101
+
102
+ - **Settings:**
103
+
104
+ - [ ] In interactive mode run `/settings` and make modifications
105
+ - [ ] Validate that setting is changed
106
+
107
+ - **Function calling:**
108
+ - [ ] In interactive mode, ask gemini to "create a file named hello.md with
109
+ the content 'hello world'" and verify the file is created correctly.
110
+
111
+ If any of these CUJs fail, the release is a no-go until a patch is applied to
112
+ the `preview` channel.
113
+
114
+ ### 3. Pre-Launch bug bash (tier 1 and 2 launches)
115
+
116
+ For high-impact releases, an organized bug bash is required to ensure a higher
117
+ level of quality and to catch issues across a wider range of environments and
118
+ use cases.
119
+
120
+ **Definition of tiers:**
121
+
122
+ - **Tier 1:** Industry-Moving News 🚀
123
+ - **Tier 2:** Important News for Our Users 📣
124
+ - **Tier 3:** Relevant, but Not Life-Changing 💡
125
+ - **Tier 4:** Bug Fixes ⚒️
126
+
127
+ **Requirement:**
128
+
129
+ A bug bash must be scheduled at least **72 hours in advance** of any Tier 1 or
130
+ Tier 2 launch.
131
+
132
+ **Rule of thumb:**
133
+
134
+ A bug bash should be considered for any release that involves:
135
+
136
+ - A blog post
137
+ - Coordinated social media announcements
138
+ - Media relations or press outreach
139
+ - A "Turbo" launch event
140
+
141
+ ## Level 3: Telemetry and data review
142
+
143
+ ### Dashboard health
144
+
145
+ - [ ] Go to `go/gemini-cli-dash`.
146
+ - [ ] Navigate to the "Tool Call" tab.
147
+ - [ ] Validate that there are no spikes in errors for the release you would like
148
+ to promote.
149
+
150
+ ### Model evaluation
151
+
152
+ - [ ] Navigate to `go/gemini-cli-offline-evals-dash`.
153
+ - [ ] Make sure that the release you want to promote's recurring run is within
154
+ average eval runs.
155
+
156
+ ## The "go/no-go" decision
157
+
158
+ Before triggering the `Release: Promote` workflow to move `preview` to `stable`:
159
+
160
+ 1. [ ] **Level 1:** CI and E2E workflows are green for the commit corresponding
161
+ to the current `preview` tag.
162
+ 2. [ ] **Level 2:** The `preview` version has been out for one week, and the
163
+ CUJ checklist has been completed successfully by a release manager. No
164
+ blocking issues have been reported.
165
+ 3. [ ] **Level 3:** Dashboard Health and Model Evaluation checks have been
166
+ completed and show no regressions.
167
+
168
+ If all checks pass, proceed with the promotion.
docs/releases.md ADDED
@@ -0,0 +1,550 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Gemini CLI releases
2
+
3
+ <!-- prettier-ignore -->
4
+ > [!IMPORTANT]
5
+ > **Coordinate with the Release Manager:** The release manager is responsible for coordinating patches and releases. Please update them before performing any of the release actions described in this document.
6
+
7
+ ## `dev` vs `prod` environment
8
+
9
+ Our release flows support both `dev` and `prod` environments.
10
+
11
+ The `dev` environment pushes to a private GitHub-hosted NPM repository, with the
12
+ package names beginning with `@google-gemini/**` instead of `@google/**`.
13
+
14
+ The `prod` environment pushes to the public global NPM registry via Wombat
15
+ Dressing Room, which is Google's system for managing NPM packages in the
16
+ `@google/**` namespace. The packages are all named `@google/**`.
17
+
18
+ More information can be found about these systems in the
19
+ [NPM Package Overview](npm.md)
20
+
21
+ ### Package scopes
22
+
23
+ | Package | `prod` (Wombat Dressing Room) | `dev` (GitHub Private NPM Repo) |
24
+ | ---------- | ----------------------------- | ----------------------------------------- |
25
+ | CLI | @google/gemini-cli | @google-gemini/gemini-cli |
26
+ | Core | @google/gemini-cli-core | @google-gemini/gemini-cli-core A2A Server |
27
+ | A2A Server | @google/gemini-cli-a2a-server | @google-gemini/gemini-cli-a2a-server |
28
+
29
+ ## Release cadence and tags
30
+
31
+ We will follow https://semver.org/ as closely as possible but will call out when
32
+ or if we have to deviate from it. Our weekly releases will be minor version
33
+ increments and any bug or hotfixes between releases will go out as patch
34
+ versions on the most recent release.
35
+
36
+ Each Tuesday ~20:00 UTC new Stable and Preview releases will be cut. The
37
+ promotion flow is:
38
+
39
+ - Code is committed to main and pushed each night to nightly
40
+ - After no more than 1 week on main, code is promoted to the `preview` channel
41
+ - After 1 week the most recent `preview` channel is promoted to `stable` channel
42
+ - Patch fixes will be produced against both `preview` and `stable` as needed,
43
+ with the final 'patch' version number incrementing each time.
44
+
45
+ ### Preview
46
+
47
+ These releases will not have been fully vetted and may contain regressions or
48
+ other outstanding issues. Help us test and install with `preview` tag.
49
+
50
+ ```bash
51
+ npm install -g @google/gemini-cli@preview
52
+ ```
53
+
54
+ ### Stable
55
+
56
+ This will be the full promotion of last week's release + any bug fixes and
57
+ validations. Use `latest` tag.
58
+
59
+ ```bash
60
+ npm install -g @google/gemini-cli@latest
61
+ ```
62
+
63
+ ### Nightly
64
+
65
+ - New releases will be published each day at UTC 00:00. This will be all changes
66
+ from the main branch as represented at time of release. It should be assumed
67
+ there are pending validations and issues. Use `nightly` tag.
68
+
69
+ ```bash
70
+ npm install -g @google/gemini-cli@nightly
71
+ ```
72
+
73
+ ## Weekly release promotion
74
+
75
+ Each Tuesday, the on-call engineer will trigger the "Promote Release" workflow.
76
+ This single action automates the entire weekly release process:
77
+
78
+ 1. **Promotes preview to stable:** The workflow identifies the latest `preview`
79
+ release and promotes it to `stable`. This becomes the new `latest` version
80
+ on npm.
81
+ 2. **Promotes nightly to preview:** The latest `nightly` release is then
82
+ promoted to become the new `preview` version.
83
+ 3. **Prepares for next nightly:** A pull request is automatically created and
84
+ merged to bump the version in `main` in preparation for the next nightly
85
+ release.
86
+
87
+ This process ensures a consistent and reliable release cadence with minimal
88
+ manual intervention.
89
+
90
+ ### Source of truth for versioning
91
+
92
+ To ensure the highest reliability, the release promotion process uses the **NPM
93
+ registry as the single source of truth** for determining the current version of
94
+ each release channel (`stable`, `preview`, and `nightly`).
95
+
96
+ 1. **Fetch from NPM:** The workflow begins by querying NPM's `dist-tags`
97
+ (`latest`, `preview`, `nightly`) to get the exact version strings for the
98
+ packages currently available to users.
99
+ 2. **Cross-check for integrity:** For each version retrieved from NPM, the
100
+ workflow performs a critical integrity check:
101
+ - It verifies that a corresponding **git tag** exists in the repository.
102
+ - It verifies that a corresponding **GitHub release** has been created.
103
+ 3. **Halt on discrepancy:** If either the git tag or the GitHub Release is
104
+ missing for a version listed on NPM, the workflow will immediately fail.
105
+ This strict check prevents promotions from a broken or incomplete previous
106
+ release and alerts the on-call engineer to a release state inconsistency
107
+ that must be manually resolved.
108
+ 4. **Calculate next version:** Only after these checks pass does the workflow
109
+ proceed to calculate the next semantic version based on the trusted version
110
+ numbers retrieved from NPM.
111
+
112
+ This NPM-first approach, backed by integrity checks, makes the release process
113
+ highly robust and prevents the kinds of versioning discrepancies that can arise
114
+ from relying solely on git history or API outputs.
115
+
116
+ ## Manual releases
117
+
118
+ For situations requiring a release outside of the regular nightly and weekly
119
+ promotion schedule, and NOT already covered by patching process, you can use the
120
+ `Release: Manual` workflow. This workflow provides a direct way to publish a
121
+ specific version from any branch, tag, or commit SHA.
122
+
123
+ ### How to create a manual release
124
+
125
+ 1. Navigate to the **Actions** tab of the repository.
126
+ 2. Select the **Release: Manual** workflow from the list.
127
+ 3. Click the **Run workflow** dropdown button.
128
+ 4. Fill in the required inputs:
129
+ - **Version**: The exact version to release (for example, `v0.6.1`). This
130
+ must be a valid semantic version with a `v` prefix.
131
+ - **Ref**: The branch, tag, or full commit SHA to release from.
132
+ - **NPM Channel**: The npm channel to publish to. The options are `preview`,
133
+ `nightly`, `latest` (for stable releases), and `dev`. The default is
134
+ `dev`.
135
+ - **Dry Run**: Leave as `true` to run all steps without publishing, or set
136
+ to `false` to perform a live release.
137
+ - **Force Skip Tests**: Set to `true` to skip the test suite. This is not
138
+ recommended for production releases.
139
+ - **Skip GitHub Release**: Set to `true` to skip creating a GitHub release
140
+ and create an npm release only.
141
+ - **Environment**: Select the appropriate environment. The `dev` environment
142
+ is intended for testing. The `prod` environment is intended for production
143
+ releases. `prod` is the default and will require authorization from a
144
+ release administrator.
145
+ 5. Click **Run workflow**.
146
+
147
+ The workflow will then proceed to test (if not skipped), build, and publish the
148
+ release. If the workflow fails during a non-dry run, it will automatically
149
+ create a GitHub issue with the failure details.
150
+
151
+ ## Rollback/rollforward
152
+
153
+ In the event that a release has a critical regression, you can quickly roll back
154
+ to a previous stable version or roll forward to a new patch by changing the npm
155
+ `dist-tag`. The `Release: Change Tags` workflow provides a safe and controlled
156
+ way to do this.
157
+
158
+ This is the preferred method for both rollbacks and rollforwards, as it does not
159
+ require a full release cycle.
160
+
161
+ ### How to change a release tag
162
+
163
+ 1. Navigate to the **Actions** tab of the repository.
164
+ 2. Select the **Release: Change Tags** workflow from the list.
165
+ 3. Click the **Run workflow** dropdown button.
166
+ 4. Fill in the required inputs:
167
+ - **Version**: The existing package version that you want to point the tag
168
+ to (for example, `0.5.0-preview-2`). This version **must** already be
169
+ published to the npm registry.
170
+ - **Channel**: The npm `dist-tag` to apply (for example, `preview`,
171
+ `stable`).
172
+ - **Dry Run**: Leave as `true` to log the action without making changes, or
173
+ set to `false` to perform the live tag change.
174
+ - **Environment**: Select the appropriate environment. The `dev` environment
175
+ is intended for testing. The `prod` environment is intended for production
176
+ releases. `prod` is the default and will require authorization from a
177
+ release administrator.
178
+ 5. Click **Run workflow**.
179
+
180
+ The workflow will then run `npm dist-tag add` for the appropriate `gemini-cli`,
181
+ `gemini-cli-core` and `gemini-cli-a2a-server` packages, pointing the specified
182
+ channel to the specified version.
183
+
184
+ ## Patching
185
+
186
+ If a critical bug that is already fixed on `main` needs to be patched on a
187
+ `stable` or `preview` release, the process is now highly automated.
188
+
189
+ ### How to patch
190
+
191
+ #### 1. Create the patch pull request
192
+
193
+ There are two ways to create a patch pull request:
194
+
195
+ **Option A: From a GitHub comment (recommended)**
196
+
197
+ After a pull request containing the fix has been merged, a maintainer can add a
198
+ comment on that same PR with the following format:
199
+
200
+ `/patch [channel]`
201
+
202
+ - **channel** (optional):
203
+ - _no channel_ - patches both stable and preview channels (default,
204
+ recommended for most fixes)
205
+ - `both` - patches both stable and preview channels (same as default)
206
+ - `stable` - patches only the stable channel
207
+ - `preview` - patches only the preview channel
208
+
209
+ Examples:
210
+
211
+ - `/patch` (patches both stable and preview - default)
212
+ - `/patch both` (patches both stable and preview - explicit)
213
+ - `/patch stable` (patches only stable)
214
+ - `/patch preview` (patches only preview)
215
+
216
+ The `Release: Patch from Comment` workflow will automatically find the merge
217
+ commit SHA and trigger the `Release: Patch (1) Create PR` workflow. If the PR is
218
+ not yet merged, it will post a comment indicating the failure.
219
+
220
+ **Option B: Manually triggering the workflow**
221
+
222
+ Navigate to the **Actions** tab and run the **Release: Patch (1) Create PR**
223
+ workflow.
224
+
225
+ - **Commit**: The full SHA of the commit on `main` that you want to cherry-pick.
226
+ - **Channel**: The channel you want to patch (`stable` or `preview`).
227
+
228
+ This workflow will automatically:
229
+
230
+ 1. Find the latest release tag for the channel.
231
+ 2. Create a release branch from that tag if one doesn't exist (for example,
232
+ `release/v0.5.1-pr-12345`).
233
+ 3. Create a new hotfix branch from the release branch.
234
+ 4. Cherry-pick your specified commit into the hotfix branch.
235
+ 5. Create a pull request from the hotfix branch back to the release branch.
236
+
237
+ #### 2. Review and merge
238
+
239
+ Review the automatically created pull request(s) to ensure the cherry-pick was
240
+ successful and the changes are correct. Once approved, merge the pull request.
241
+
242
+ <!-- prettier-ignore -->
243
+ > [!WARNING]
244
+ > The `release/*` branches are protected by branch protection
245
+ > rules. A pull request to one of these branches requires at least one review from
246
+ > a code owner before it can be merged. This ensures that no unauthorized code is
247
+ > released.
248
+
249
+ #### 2.5. Adding multiple commits to a hotfix (advanced)
250
+
251
+ If you need to include multiple fixes in a single patch release, you can add
252
+ additional commits to the hotfix branch after the initial patch PR has been
253
+ created:
254
+
255
+ 1. **Start with the primary fix**: Use `/patch` (or `/patch both`) on the most
256
+ important PR to create the initial hotfix branch and PR.
257
+
258
+ 2. **Checkout the hotfix branch locally**:
259
+
260
+ ```bash
261
+ git fetch origin
262
+ git checkout hotfix/v0.5.1/stable/cherry-pick-abc1234 # Use the actual branch name from the PR
263
+ ```
264
+
265
+ 3. **Cherry-pick additional commits**:
266
+
267
+ ```bash
268
+ git cherry-pick <commit-sha-1>
269
+ git cherry-pick <commit-sha-2>
270
+ # Add as many commits as needed
271
+ ```
272
+
273
+ 4. **Push the updated branch**:
274
+
275
+ ```bash
276
+ git push origin hotfix/v0.5.1/stable/cherry-pick-abc1234
277
+ ```
278
+
279
+ 5. **Test and review**: The existing patch PR will automatically update with
280
+ your additional commits. Test thoroughly since you're now releasing multiple
281
+ changes together.
282
+
283
+ 6. **Update the PR description**: Consider updating the PR title and description
284
+ to reflect that it includes multiple fixes.
285
+
286
+ This approach lets you group related fixes into a single patch release while
287
+ maintaining full control over what gets included and how conflicts are resolved.
288
+
289
+ #### 3. Automatic release
290
+
291
+ Upon merging the pull request, the `Release: Patch (2) Trigger` workflow is
292
+ automatically triggered. It will then start the `Release: Patch (3) Release`
293
+ workflow, which will:
294
+
295
+ 1. Build and test the patched code.
296
+ 2. Publish the new patch version to npm.
297
+ 3. Create a new GitHub release with the patch notes.
298
+
299
+ This fully automated process ensures that patches are created and released
300
+ consistently and reliably.
301
+
302
+ #### Troubleshooting: Older branch workflows
303
+
304
+ **Issue**: If the patch trigger workflow fails with errors like "Resource not
305
+ accessible by integration" or references to non-existent workflow files (for
306
+ example, `patch-release.yml`), this indicates the hotfix branch contains an
307
+ outdated version of the workflow files.
308
+
309
+ **Root cause**: When a PR is merged, GitHub Actions runs the workflow definition
310
+ from the **source branch** (the hotfix branch), not from the target branch (the
311
+ release branch). If the hotfix branch was created from an older release branch
312
+ that predates workflow improvements, it will use the old workflow logic.
313
+
314
+ **Solutions**:
315
+
316
+ **Option 1: Manual trigger (quick fix)** Manually trigger the updated workflow
317
+ from the branch with the latest workflow code:
318
+
319
+ ```bash
320
+ # For a preview channel patch with tests skipped
321
+ gh workflow run release-patch-2-trigger.yml --ref <branch-with-updated-workflow> \
322
+ --field ref="hotfix/v0.6.0-preview.2/preview/cherry-pick-abc1234" \
323
+ --field workflow_ref=<branch-with-updated-workflow> \
324
+ --field dry_run=false \
325
+ --field force_skip_tests=true
326
+
327
+ # For a stable channel patch
328
+ gh workflow run release-patch-2-trigger.yml --ref <branch-with-updated-workflow> \
329
+ --field ref="hotfix/v0.5.1/stable/cherry-pick-abc1234" \
330
+ --field workflow_ref=<branch-with-updated-workflow> \
331
+ --field dry_run=false \
332
+ --field force_skip_tests=false
333
+
334
+ # Example using main branch (most common case)
335
+ gh workflow run release-patch-2-trigger.yml --ref main \
336
+ --field ref="hotfix/v0.6.0-preview.2/preview/cherry-pick-abc1234" \
337
+ --field workflow_ref=main \
338
+ --field dry_run=false \
339
+ --field force_skip_tests=true
340
+ ```
341
+
342
+ **Note**: Replace `<branch-with-updated-workflow>` with the branch containing
343
+ the latest workflow improvements (usually `main`, but could be a feature branch
344
+ if testing updates).
345
+
346
+ **Option 2: Update the hotfix branch** Merge the latest main branch into your
347
+ hotfix branch to get the updated workflows:
348
+
349
+ ```bash
350
+ git checkout hotfix/v0.6.0-preview.2/preview/cherry-pick-abc1234
351
+ git merge main
352
+ git push
353
+ ```
354
+
355
+ Then close and reopen the PR to retrigger the workflow with the updated version.
356
+
357
+ **Option 3: Direct release trigger** Skip the trigger workflow entirely and
358
+ directly run the release workflow:
359
+
360
+ ```bash
361
+ # Replace channel and release_ref with appropriate values
362
+ gh workflow run release-patch-3-release.yml --ref main \
363
+ --field type="preview" \
364
+ --field dry_run=false \
365
+ --field force_skip_tests=true \
366
+ --field release_ref="release/v0.6.0-preview.2"
367
+ ```
368
+
369
+ ### Docker
370
+
371
+ We also run a Google cloud build called
372
+ [release-docker.yml](../.gcp/release-docker.yml). Which publishes the sandbox
373
+ docker to match your release. This will also be moved to GH and combined with
374
+ the main release file once service account permissions are sorted out.
375
+
376
+ ## Release validation
377
+
378
+ After pushing a new release smoke testing should be performed to ensure that the
379
+ packages are working as expected. This can be done by installing the packages
380
+ locally and running a set of tests to ensure that they are functioning
381
+ correctly.
382
+
383
+ - `npx -y @google/gemini-cli@latest --version` to validate the push worked as
384
+ expected if you were not doing a rc or dev tag
385
+ - `npx -y @google/gemini-cli@<release tag> --version` to validate the tag pushed
386
+ appropriately
387
+ - _This is destructive locally_
388
+ `npm uninstall @google/gemini-cli && npm uninstall -g @google/gemini-cli && npm cache clean --force && npm install @google/gemini-cli@<version>`
389
+ - Smoke testing a basic run through of exercising a few llm commands and tools
390
+ is recommended to ensure that the packages are working as expected. We'll
391
+ codify this more in the future.
392
+
393
+ ## Local testing and validation: Changes to the packaging and publishing process
394
+
395
+ If you need to test the release process without actually publishing to NPM or
396
+ creating a public GitHub release, you can trigger the workflow manually from the
397
+ GitHub UI.
398
+
399
+ 1. Go to the
400
+ [Actions tab](https://github.com/google-gemini/gemini-cli/actions/workflows/release-manual.yml)
401
+ of the repository.
402
+ 2. Click on the "Run workflow" dropdown.
403
+ 3. Leave the `dry_run` option checked (`true`).
404
+ 4. Click the "Run workflow" button.
405
+
406
+ This will run the entire release process but will skip the `npm publish` and
407
+ `gh release create` steps. You can inspect the workflow logs to ensure
408
+ everything is working as expected.
409
+
410
+ It is crucial to test any changes to the packaging and publishing process
411
+ locally before committing them. This ensures that the packages will be published
412
+ correctly and that they will work as expected when installed by a user.
413
+
414
+ To validate your changes, you can perform a dry run of the publishing process.
415
+ This will simulate the publishing process without actually publishing the
416
+ packages to the npm registry.
417
+
418
+ ```bash
419
+ npm_package_version=9.9.9 SANDBOX_IMAGE_REGISTRY="registry" SANDBOX_IMAGE_NAME="thename" npm run publish:npm --dry-run
420
+ ```
421
+
422
+ This command will do the following:
423
+
424
+ 1. Build all the packages.
425
+ 2. Run all the prepublish scripts.
426
+ 3. Create the package tarballs that would be published to npm.
427
+ 4. Print a summary of the packages that would be published.
428
+
429
+ You can then inspect the generated tarballs to ensure that they contain the
430
+ correct files and that the `package.json` files have been updated correctly. The
431
+ tarballs will be created in the root of each package's directory (for example,
432
+ `packages/cli/google-gemini-cli-0.1.6.tgz`).
433
+
434
+ By performing a dry run, you can be confident that your changes to the packaging
435
+ process are correct and that the packages will be published successfully.
436
+
437
+ ## Release deep dive
438
+
439
+ The release process creates two distinct types of artifacts for different
440
+ distribution channels: standard packages for the NPM registry and a single,
441
+ self-contained executable for GitHub Releases.
442
+
443
+ Here are the key stages:
444
+
445
+ **Stage 1: Pre-release sanity checks and versioning**
446
+
447
+ - **What happens:** Before any files are moved, the process ensures the project
448
+ is in a good state. This involves running tests, linting, and type-checking
449
+ (`npm run preflight`). The version number in the root `package.json` and
450
+ `packages/cli/package.json` is updated to the new release version.
451
+
452
+ **Stage 2: Building the source code for NPM**
453
+
454
+ - **What happens:** The TypeScript source code in `packages/core/src` and
455
+ `packages/cli/src` is compiled into standard JavaScript.
456
+ - **File movement:**
457
+ - `packages/core/src/**/*.ts` -> compiled to -> `packages/core/dist/`
458
+ - `packages/cli/src/**/*.ts` -> compiled to -> `packages/cli/dist/`
459
+ - **Why:** The TypeScript code written during development needs to be converted
460
+ into plain JavaScript that can be run by Node.js. The `core` package is built
461
+ first as the `cli` package depends on it.
462
+
463
+ **Stage 3: Publishing standard packages to NPM**
464
+
465
+ - **What happens:** The `npm publish` command is run for the
466
+ `@google/gemini-cli-core` and `@google/gemini-cli` packages.
467
+ - **Why:** This publishes them as standard Node.js packages. Users installing
468
+ via `npm install -g @google/gemini-cli` will download these packages, and
469
+ `npm` will handle installing the `@google/gemini-cli-core` dependency
470
+ automatically. The code in these packages is not bundled into a single file.
471
+
472
+ **Stage 4: Assembling and creating the GitHub release asset**
473
+
474
+ This stage happens _after_ the NPM publish and creates the single-file
475
+ executable that enables `npx` usage directly from the GitHub repository.
476
+
477
+ 1. **The JavaScript bundle is created:**
478
+
479
+ - **What happens:** The built JavaScript from both `packages/core/dist` and
480
+ `packages/cli/dist`, along with all third-party JavaScript dependencies,
481
+ are bundled by `esbuild` into a single, executable JavaScript file (for
482
+ example, `gemini.js`). The `node-pty` library is excluded from this bundle
483
+ as it contains native binaries.
484
+ - **Why:** This creates a single, optimized file that contains all the
485
+ necessary application code. It simplifies execution for users who want to
486
+ run the CLI without a full `npm install`, as all dependencies (including
487
+ the `core` package) are included directly.
488
+
489
+ 2. **The `bundle` directory is assembled:**
490
+
491
+ - **What happens:** A temporary `bundle` folder is created at the project
492
+ root. The single `gemini.js` executable is placed inside it, along with
493
+ other essential files.
494
+ - **File movement:**
495
+ - `gemini.js` (from esbuild) -> `bundle/gemini.js`
496
+ - `README.md` -> `bundle/README.md`
497
+ - `LICENSE` -> `bundle/LICENSE`
498
+ - `packages/cli/src/utils/*.sb` (sandbox profiles) -> `bundle/`
499
+ - **Why:** This creates a clean, self-contained directory with everything
500
+ needed to run the CLI and understand its license and usage.
501
+
502
+ 3. **The GitHub release is created:**
503
+ - **What happens:** The contents of the `bundle` directory, including the
504
+ `gemini.js` executable, are attached as assets to a new GitHub Release.
505
+ - **Why:** This makes the single-file version of the CLI available for
506
+ direct download and enables the
507
+ `npx https://github.com/google-gemini/gemini-cli` command, which downloads
508
+ and runs this specific bundled asset.
509
+
510
+ **Summary of artifacts**
511
+
512
+ - **NPM:** Publishes standard, un-bundled Node.js packages. The primary artifact
513
+ is the code in `packages/cli/dist`, which depends on
514
+ `@google/gemini-cli-core`.
515
+ - **GitHub release:** Publishes a single, bundled `gemini.js` file that contains
516
+ all dependencies, for easy execution via `npx`.
517
+
518
+ This dual-artifact process ensures that both traditional `npm` users and those
519
+ who prefer the convenience of `npx` have an optimized experience.
520
+
521
+ ## Notifications
522
+
523
+ Failing release workflows will automatically create an issue with the label
524
+ `release-failure`.
525
+
526
+ A notification will be posted to the maintainer's chat channel when issues with
527
+ this type are created.
528
+
529
+ ### Modifying chat notifications
530
+
531
+ Notifications use
532
+ [GitHub for Google Chat](https://workspace.google.com/marketplace/app/github_for_google_chat/536184076190).
533
+ To modify the notifications, use `/github-settings` within the chat space.
534
+
535
+ <!-- prettier-ignore -->
536
+ > [!WARNING]
537
+ > The following instructions describe a fragile workaround that depends on the
538
+ > internal structure of the chat application's UI. It is likely to break with
539
+ > future updates.
540
+
541
+ The list of available labels is not currently populated correctly. If you want
542
+ to add a label that does not appear alphabetically in the first 30 labels in the
543
+ repo, you must use your browser's developer tools to manually modify the UI:
544
+
545
+ 1. Open your browser's developer tools (for example, Chrome DevTools).
546
+ 2. In the `/github-settings` dialog, inspect the list of labels.
547
+ 3. Locate one of the `<li>` elements representing a label.
548
+ 4. In the HTML, modify the `data-option-value` attribute of that `<li>` element
549
+ to the desired label name (for example, `release-failure`).
550
+ 5. Click on your modified label in the UI to select it, then save your settings.
docs/sidebar.json ADDED
@@ -0,0 +1,298 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "label": "docs_tab",
4
+ "items": [
5
+ {
6
+ "label": "Get started",
7
+ "items": [
8
+ { "label": "Overview", "slug": "docs" },
9
+ { "label": "Quickstart", "slug": "docs/get-started" },
10
+ { "label": "Installation", "slug": "docs/get-started/installation" },
11
+ {
12
+ "label": "Authentication",
13
+ "slug": "docs/get-started/authentication"
14
+ },
15
+ { "label": "CLI cheatsheet", "slug": "docs/cli/cli-reference" },
16
+ {
17
+ "label": "Gemini 3 on Gemini CLI",
18
+ "slug": "docs/get-started/gemini-3"
19
+ }
20
+ ]
21
+ },
22
+ {
23
+ "label": "Use Gemini CLI",
24
+ "items": [
25
+ {
26
+ "label": "File management",
27
+ "slug": "docs/cli/tutorials/file-management"
28
+ },
29
+ {
30
+ "label": "Get started with Agent Skills",
31
+ "slug": "docs/cli/tutorials/skills-getting-started"
32
+ },
33
+ {
34
+ "label": "Manage context and memory",
35
+ "slug": "docs/cli/tutorials/memory-management"
36
+ },
37
+ {
38
+ "label": "Execute shell commands",
39
+ "slug": "docs/cli/tutorials/shell-commands"
40
+ },
41
+ {
42
+ "label": "Manage sessions and history",
43
+ "slug": "docs/cli/tutorials/session-management"
44
+ },
45
+ {
46
+ "label": "Plan tasks with todos",
47
+ "slug": "docs/cli/tutorials/task-planning"
48
+ },
49
+ {
50
+ "label": "Use Plan Mode with model steering",
51
+ "badge": "🔬",
52
+ "slug": "docs/cli/tutorials/plan-mode-steering"
53
+ },
54
+ {
55
+ "label": "Web search and fetch",
56
+ "slug": "docs/cli/tutorials/web-tools"
57
+ },
58
+ {
59
+ "label": "Set up an MCP server",
60
+ "slug": "docs/cli/tutorials/mcp-setup"
61
+ },
62
+ { "label": "Automate tasks", "slug": "docs/cli/tutorials/automation" }
63
+ ]
64
+ },
65
+ {
66
+ "label": "Features",
67
+ "items": [
68
+ {
69
+ "label": "Extensions",
70
+ "collapsed": true,
71
+ "items": [
72
+ {
73
+ "label": "Overview",
74
+ "slug": "docs/extensions"
75
+ },
76
+ {
77
+ "label": "User guide: Install and manage",
78
+ "link": "/docs/extensions/#manage-extensions"
79
+ },
80
+ {
81
+ "label": "Developer guide: Build extensions",
82
+ "slug": "docs/extensions/writing-extensions"
83
+ },
84
+ {
85
+ "label": "Developer guide: Best practices",
86
+ "slug": "docs/extensions/best-practices"
87
+ },
88
+ {
89
+ "label": "Developer guide: Releasing",
90
+ "slug": "docs/extensions/releasing"
91
+ },
92
+ {
93
+ "label": "Developer guide: Reference",
94
+ "slug": "docs/extensions/reference"
95
+ }
96
+ ]
97
+ },
98
+ {
99
+ "label": "Agent Skills",
100
+ "collapsed": true,
101
+ "items": [
102
+ { "label": "Overview", "slug": "docs/cli/skills" },
103
+ {
104
+ "label": "Get started with Agent Skills",
105
+ "slug": "docs/cli/tutorials/skills-getting-started"
106
+ },
107
+ {
108
+ "label": "Creating Agent Skills",
109
+ "slug": "docs/cli/creating-skills"
110
+ },
111
+ {
112
+ "label": "Using Agent Skills",
113
+ "slug": "docs/cli/using-agent-skills"
114
+ },
115
+ {
116
+ "label": "Developer guide: Best practices",
117
+ "slug": "docs/cli/skills-best-practices"
118
+ }
119
+ ]
120
+ },
121
+ {
122
+ "label": "Auto Memory",
123
+ "badge": "🔬",
124
+ "slug": "docs/cli/auto-memory"
125
+ },
126
+ { "label": "Checkpointing", "slug": "docs/cli/checkpointing" },
127
+ { "label": "Headless mode", "slug": "docs/cli/headless" },
128
+ {
129
+ "label": "Git worktrees",
130
+ "badge": "🔬",
131
+ "slug": "docs/cli/git-worktrees"
132
+ },
133
+ {
134
+ "label": "Hooks",
135
+ "collapsed": true,
136
+ "items": [
137
+ { "label": "Overview", "slug": "docs/hooks" },
138
+ { "label": "Reference", "slug": "docs/hooks/reference" }
139
+ ]
140
+ },
141
+ {
142
+ "label": "IDE integration",
143
+ "collapsed": true,
144
+ "items": [
145
+ { "label": "Overview", "slug": "docs/ide-integration" },
146
+ {
147
+ "label": "Developer guide: ACP mode",
148
+ "slug": "docs/cli/acp-mode"
149
+ }
150
+ ]
151
+ },
152
+ {
153
+ "label": "MCP servers",
154
+ "collapsed": true,
155
+ "items": [
156
+ { "label": "Overview", "slug": "docs/tools/mcp-server" },
157
+ { "label": "Resource tools", "slug": "docs/tools/mcp-resources" }
158
+ ]
159
+ },
160
+ { "label": "Model routing", "slug": "docs/cli/model-routing" },
161
+ { "label": "Model selection", "slug": "docs/cli/model" },
162
+ {
163
+ "label": "Model steering",
164
+ "badge": "🔬",
165
+ "slug": "docs/cli/model-steering"
166
+ },
167
+ {
168
+ "label": "Notifications",
169
+ "badge": "🔬",
170
+ "slug": "docs/cli/notifications"
171
+ },
172
+ { "label": "Plan mode", "slug": "docs/cli/plan-mode" },
173
+ {
174
+ "label": "Subagents",
175
+ "slug": "docs/core/subagents"
176
+ },
177
+ {
178
+ "label": "Remote subagents",
179
+ "slug": "docs/core/remote-agents"
180
+ },
181
+ { "label": "Rewind", "slug": "docs/cli/rewind" },
182
+ { "label": "Sandboxing", "slug": "docs/cli/sandbox" },
183
+ { "label": "Settings", "slug": "docs/cli/settings" },
184
+ { "label": "Telemetry", "slug": "docs/cli/telemetry" },
185
+ { "label": "Token caching", "slug": "docs/cli/token-caching" }
186
+ ]
187
+ },
188
+ {
189
+ "label": "Configuration",
190
+ "items": [
191
+ { "label": "Custom commands", "slug": "docs/cli/custom-commands" },
192
+ {
193
+ "label": "Enterprise configuration",
194
+ "slug": "docs/cli/enterprise"
195
+ },
196
+ {
197
+ "label": "Ignore files (.geminiignore)",
198
+ "slug": "docs/cli/gemini-ignore"
199
+ },
200
+ {
201
+ "label": "Model configuration",
202
+ "slug": "docs/cli/generation-settings"
203
+ },
204
+ {
205
+ "label": "Project context (GEMINI.md)",
206
+ "slug": "docs/cli/gemini-md"
207
+ },
208
+ { "label": "Settings", "slug": "docs/cli/settings" },
209
+ {
210
+ "label": "System prompt override",
211
+ "slug": "docs/cli/system-prompt"
212
+ },
213
+ { "label": "Themes", "slug": "docs/cli/themes" },
214
+ { "label": "Trusted folders", "slug": "docs/cli/trusted-folders" }
215
+ ]
216
+ },
217
+ {
218
+ "label": "Development",
219
+ "items": [
220
+ {
221
+ "label": "Behavioral evaluations",
222
+ "slug": "docs/behavioral-evals"
223
+ },
224
+ { "label": "Contribution guide", "slug": "docs/contributing" },
225
+ { "label": "Integration testing", "slug": "docs/integration-tests" },
226
+ {
227
+ "label": "Issue and PR automation",
228
+ "slug": "docs/issue-and-pr-automation"
229
+ },
230
+ { "label": "Local development", "slug": "docs/local-development" },
231
+ { "label": "NPM package structure", "slug": "docs/npm" }
232
+ ]
233
+ }
234
+ ]
235
+ },
236
+ {
237
+ "label": "reference_tab",
238
+ "items": [
239
+ {
240
+ "label": "Reference",
241
+ "items": [
242
+ { "label": "Command reference", "slug": "docs/reference/commands" },
243
+ {
244
+ "label": "Configuration reference",
245
+ "slug": "docs/reference/configuration"
246
+ },
247
+ {
248
+ "label": "Keyboard shortcuts",
249
+ "slug": "docs/reference/keyboard-shortcuts"
250
+ },
251
+ {
252
+ "label": "Memory import processor",
253
+ "slug": "docs/reference/memport"
254
+ },
255
+ { "label": "Policy engine", "slug": "docs/reference/policy-engine" },
256
+ { "label": "Tools reference", "slug": "docs/reference/tools" }
257
+ ]
258
+ }
259
+ ]
260
+ },
261
+ {
262
+ "label": "resources_tab",
263
+ "items": [
264
+ {
265
+ "label": "Resources",
266
+ "items": [
267
+ { "label": "FAQ", "slug": "docs/resources/faq" },
268
+ {
269
+ "label": "Quota and pricing",
270
+ "slug": "docs/resources/quota-and-pricing"
271
+ },
272
+ {
273
+ "label": "Terms and privacy",
274
+ "slug": "docs/resources/tos-privacy"
275
+ },
276
+ {
277
+ "label": "Troubleshooting",
278
+ "slug": "docs/resources/troubleshooting"
279
+ },
280
+ { "label": "Uninstall", "slug": "docs/resources/uninstall" }
281
+ ]
282
+ }
283
+ ]
284
+ },
285
+ {
286
+ "label": "releases_tab",
287
+ "items": [
288
+ {
289
+ "label": "Releases",
290
+ "items": [
291
+ { "label": "Release notes", "slug": "docs/changelogs/" },
292
+ { "label": "Stable release", "slug": "docs/changelogs/latest" },
293
+ { "label": "Preview release", "slug": "docs/changelogs/preview" }
294
+ ]
295
+ }
296
+ ]
297
+ }
298
+ ]
integration-tests/acp-env-auth.test.ts ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig } from './test-helper.js';
9
+ import { spawn, ChildProcess } from 'node:child_process';
10
+ import { join, resolve } from 'node:path';
11
+ import { writeFileSync, mkdirSync } from 'node:fs';
12
+ import { Writable, Readable } from 'node:stream';
13
+ import { env } from 'node:process';
14
+ import * as acp from '@agentclientprotocol/sdk';
15
+
16
+ const sandboxEnv = env['GEMINI_SANDBOX'];
17
+ const itMaybe = sandboxEnv && sandboxEnv !== 'false' ? it.skip : it;
18
+
19
+ class MockClient implements acp.Client {
20
+ updates: acp.SessionNotification[] = [];
21
+ sessionUpdate = async (params: acp.SessionNotification) => {
22
+ this.updates.push(params);
23
+ };
24
+ requestPermission = async (): Promise<acp.RequestPermissionResponse> => {
25
+ throw new Error('unexpected');
26
+ };
27
+ }
28
+
29
+ describe.skip('ACP Environment and Auth', () => {
30
+ let rig: TestRig;
31
+ let child: ChildProcess | undefined;
32
+
33
+ beforeEach(() => {
34
+ rig = new TestRig();
35
+ });
36
+
37
+ afterEach(async () => {
38
+ child?.kill();
39
+ child = undefined;
40
+ await rig.cleanup();
41
+ });
42
+
43
+ itMaybe(
44
+ 'should load .env from project directory and use the provided API key',
45
+ async () => {
46
+ rig.setup('acp-env-loading');
47
+
48
+ // Create a project directory with a .env file containing a recognizable invalid key
49
+ const projectDir = resolve(join(rig.testDir!, 'project'));
50
+ mkdirSync(projectDir, { recursive: true });
51
+ writeFileSync(
52
+ join(projectDir, '.env'),
53
+ 'GEMINI_API_KEY=test-key-from-env\n',
54
+ );
55
+
56
+ const bundlePath = join(import.meta.dirname, '..', 'bundle/gemini.js');
57
+
58
+ child = spawn('node', [bundlePath, '--acp'], {
59
+ cwd: rig.homeDir!,
60
+ stdio: ['pipe', 'pipe', 'inherit'],
61
+ env: {
62
+ ...process.env,
63
+ GEMINI_CLI_HOME: rig.homeDir!,
64
+ GEMINI_API_KEY: undefined,
65
+ VERBOSE: 'true',
66
+ },
67
+ });
68
+
69
+ const input = Writable.toWeb(child.stdin!);
70
+ const output = Readable.toWeb(
71
+ child.stdout!,
72
+ ) as ReadableStream<Uint8Array>;
73
+ const testClient = new MockClient();
74
+ const stream = acp.ndJsonStream(input, output);
75
+ const connection = new acp.ClientSideConnection(() => testClient, stream);
76
+
77
+ await connection.initialize({
78
+ protocolVersion: acp.PROTOCOL_VERSION,
79
+ clientCapabilities: {
80
+ fs: { readTextFile: false, writeTextFile: false },
81
+ },
82
+ });
83
+
84
+ // 1. newSession should succeed because it finds the key in .env
85
+ const { sessionId } = await connection.newSession({
86
+ cwd: projectDir,
87
+ mcpServers: [],
88
+ });
89
+
90
+ expect(sessionId).toBeDefined();
91
+
92
+ // 2. prompt should fail because the key is invalid,
93
+ // but the error should come from the API, not the internal auth check.
94
+ await expect(
95
+ connection.prompt({
96
+ sessionId,
97
+ prompt: [{ type: 'text', text: 'hello' }],
98
+ }),
99
+ ).rejects.toSatisfy((error: unknown) => {
100
+ const acpError = error as acp.RequestError;
101
+ const errorData = acpError.data as
102
+ | { error?: { message?: string } }
103
+ | undefined;
104
+ const message = String(errorData?.error?.message || acpError.message);
105
+ // It should NOT be our internal "Authentication required" message
106
+ expect(message).not.toContain('Authentication required');
107
+ // It SHOULD be an API error mentioning the invalid key
108
+ expect(message).toContain('API key not valid');
109
+ return true;
110
+ });
111
+
112
+ child.stdin!.end();
113
+ },
114
+ );
115
+
116
+ itMaybe(
117
+ 'should fail with authRequired when no API key is found',
118
+ async () => {
119
+ rig.setup('acp-auth-failure');
120
+
121
+ const bundlePath = join(import.meta.dirname, '..', 'bundle/gemini.js');
122
+
123
+ child = spawn('node', [bundlePath, '--acp'], {
124
+ cwd: rig.homeDir!,
125
+ stdio: ['pipe', 'pipe', 'inherit'],
126
+ env: {
127
+ ...process.env,
128
+ GEMINI_CLI_HOME: rig.homeDir!,
129
+ GEMINI_API_KEY: undefined,
130
+ VERBOSE: 'true',
131
+ },
132
+ });
133
+
134
+ const input = Writable.toWeb(child.stdin!);
135
+ const output = Readable.toWeb(
136
+ child.stdout!,
137
+ ) as ReadableStream<Uint8Array>;
138
+ const testClient = new MockClient();
139
+ const stream = acp.ndJsonStream(input, output);
140
+ const connection = new acp.ClientSideConnection(() => testClient, stream);
141
+
142
+ await connection.initialize({
143
+ protocolVersion: acp.PROTOCOL_VERSION,
144
+ clientCapabilities: {
145
+ fs: { readTextFile: false, writeTextFile: false },
146
+ },
147
+ });
148
+
149
+ await expect(
150
+ connection.newSession({
151
+ cwd: resolve(rig.testDir!),
152
+ mcpServers: [],
153
+ }),
154
+ ).rejects.toMatchObject({
155
+ message: expect.stringContaining(
156
+ 'Gemini API key is missing or not configured.',
157
+ ),
158
+ });
159
+
160
+ child.stdin!.end();
161
+ },
162
+ );
163
+ });
integration-tests/acp-telemetry.test.ts ADDED
@@ -0,0 +1,116 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig } from './test-helper.js';
9
+ import { spawn, ChildProcess } from 'node:child_process';
10
+ import { join } from 'node:path';
11
+ import { readFileSync, existsSync } from 'node:fs';
12
+ import { Writable, Readable } from 'node:stream';
13
+ import { env } from 'node:process';
14
+ import * as acp from '@agentclientprotocol/sdk';
15
+
16
+ // Skip in sandbox mode - test spawns CLI directly which behaves differently in containers
17
+ const sandboxEnv = env['GEMINI_SANDBOX'];
18
+ const itMaybe = sandboxEnv && sandboxEnv !== 'false' ? it.skip : it;
19
+
20
+ // Reuse existing fake responses that return a simple "Hello" response
21
+ const SIMPLE_RESPONSE_PATH = 'hooks-system.session-startup.responses';
22
+
23
+ class SessionUpdateCollector implements acp.Client {
24
+ updates: acp.SessionNotification[] = [];
25
+
26
+ sessionUpdate = async (params: acp.SessionNotification) => {
27
+ this.updates.push(params);
28
+ };
29
+
30
+ requestPermission = async (): Promise<acp.RequestPermissionResponse> => {
31
+ throw new Error('unexpected');
32
+ };
33
+ }
34
+
35
+ describe('ACP telemetry', () => {
36
+ let rig: TestRig;
37
+ let child: ChildProcess | undefined;
38
+
39
+ beforeEach(() => {
40
+ rig = new TestRig();
41
+ });
42
+
43
+ afterEach(async () => {
44
+ child?.kill();
45
+ child = undefined;
46
+ await rig.cleanup();
47
+ });
48
+
49
+ itMaybe('should flush telemetry when connection closes', async () => {
50
+ rig.setup('acp-telemetry-flush', {
51
+ fakeResponsesPath: join(import.meta.dirname, SIMPLE_RESPONSE_PATH),
52
+ });
53
+
54
+ const telemetryPath = join(rig.homeDir!, 'telemetry.log');
55
+ const bundlePath = join(import.meta.dirname, '..', 'bundle/gemini.js');
56
+
57
+ child = spawn(
58
+ 'node',
59
+ [
60
+ bundlePath,
61
+ '--acp',
62
+ '--fake-responses',
63
+ join(rig.testDir!, 'fake-responses.json'),
64
+ ],
65
+ {
66
+ cwd: rig.testDir!,
67
+ stdio: ['pipe', 'pipe', 'inherit'],
68
+ env: {
69
+ ...process.env,
70
+ GEMINI_API_KEY: 'fake-key',
71
+ GEMINI_CLI_HOME: rig.homeDir!,
72
+ GEMINI_TELEMETRY_ENABLED: 'true',
73
+ GEMINI_TELEMETRY_TRACES_ENABLED: 'true',
74
+ GEMINI_TELEMETRY_TARGET: 'local',
75
+ GEMINI_TELEMETRY_OUTFILE: telemetryPath,
76
+ },
77
+ },
78
+ );
79
+
80
+ const input = Writable.toWeb(child.stdin!);
81
+ const output = Readable.toWeb(child.stdout!) as ReadableStream<Uint8Array>;
82
+ const testClient = new SessionUpdateCollector();
83
+ const stream = acp.ndJsonStream(input, output);
84
+ const connection = new acp.ClientSideConnection(() => testClient, stream);
85
+
86
+ await connection.initialize({
87
+ protocolVersion: acp.PROTOCOL_VERSION,
88
+ clientCapabilities: { fs: { readTextFile: false, writeTextFile: false } },
89
+ });
90
+
91
+ const { sessionId } = await connection.newSession({
92
+ cwd: rig.testDir!,
93
+ mcpServers: [],
94
+ });
95
+
96
+ await connection.prompt({
97
+ sessionId,
98
+ prompt: [{ type: 'text', text: 'Say hello' }],
99
+ });
100
+
101
+ expect(JSON.stringify(testClient.updates)).toContain('Hello');
102
+
103
+ // Close stdin to trigger telemetry flush via runExitCleanup()
104
+ child.stdin!.end();
105
+ await new Promise<void>((resolve) => {
106
+ child!.on('close', () => resolve());
107
+ });
108
+ child = undefined;
109
+
110
+ // gen_ai.output.messages is the last OTEL log emitted (after prompt response)
111
+ expect(existsSync(telemetryPath)).toBe(true);
112
+ expect(readFileSync(telemetryPath, 'utf-8')).toContain(
113
+ 'gen_ai.output.messages',
114
+ );
115
+ });
116
+ });
integration-tests/api-resilience.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Part 1. "}],"role":"model"},"index":0}]},{"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":10,"totalTokenCount":110}},{"candidates":[{"content":{"parts":[{"text":"Part 2."}],"role":"model"},"index":0,"finishReason":"STOP"}]}]}
integration-tests/api-resilience.test.ts ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2026 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig } from './test-helper.js';
9
+ import { join, dirname } from 'node:path';
10
+ import { fileURLToPath } from 'node:url';
11
+
12
+ describe('API Resilience E2E', () => {
13
+ let rig: TestRig;
14
+
15
+ beforeEach(() => {
16
+ rig = new TestRig();
17
+ });
18
+
19
+ afterEach(async () => {
20
+ await rig.cleanup();
21
+ });
22
+
23
+ it('should not crash when receiving metadata-only chunks in a stream', async () => {
24
+ await rig.setup('api-resilience-metadata-only', {
25
+ fakeResponsesPath: join(
26
+ dirname(fileURLToPath(import.meta.url)),
27
+ 'api-resilience.responses',
28
+ ),
29
+ settings: {
30
+ planSettings: { modelRouting: false },
31
+ },
32
+ });
33
+
34
+ // Run the CLI with a simple prompt.
35
+ // The fake responses will provide a stream with a metadata-only chunk in the middle.
36
+ // We use gemini-3-pro-preview to minimize internal service calls.
37
+ const result = await rig.run({
38
+ args: ['hi', '--model', 'gemini-3-pro-preview'],
39
+ });
40
+
41
+ // Verify the output contains text from the normal chunks.
42
+ // If the CLI crashed on the metadata chunk, rig.run would throw.
43
+ expect(result).toContain('Part 1.');
44
+ expect(result).toContain('Part 2.');
45
+
46
+ // Verify telemetry event for the prompt was still generated
47
+ const hasUserPromptEvent = await rig.waitForTelemetryEvent('user_prompt');
48
+ expect(hasUserPromptEvent).toBe(true);
49
+ });
50
+ });
integration-tests/browser-agent-localhost.multistep.responses ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll go through the multi-step flow on the localhost server."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Navigate to http://127.0.0.1:18923/multi-step/step1.html, fill in 'testuser' as the username, click Next, then on step 2 select 'Option B' and click Finish. Report the final result page content."}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"navigate_page","args":{"url":"http://127.0.0.1:18923/multi-step/step1.html"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":20,"totalTokenCount":120}}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"fill","args":{"selector":"#username","value":"testuser"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":150,"candidatesTokenCount":25,"totalTokenCount":175}}]}
4
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"click","args":{"selector":"#next-btn"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":180,"candidatesTokenCount":20,"totalTokenCount":200}}]}
5
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":210,"candidatesTokenCount":15,"totalTokenCount":225}}]}
6
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"click","args":{"selector":"#finish-btn"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":240,"candidatesTokenCount":20,"totalTokenCount":260}}]}
7
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":270,"candidatesTokenCount":15,"totalTokenCount":285}}]}
8
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"result":{"success":true,"summary":"Completed all steps. Step 1: entered username 'testuser'. Step 2: selected default option. Final result page shows 'Multi-Step Complete' with '✓ Complete' status badge."}}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":300,"candidatesTokenCount":40,"totalTokenCount":340}}]}
9
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I've completed the multi-step flow:\n\n1. **Step 1**: Entered 'testuser' as username and clicked Next\n2. **Step 2**: Confirmed selection and clicked Finish\n3. **Result**: Final page shows 'Multi-Step Complete' with a '✓ Complete' status badge\n\nAll steps were successfully navigated."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":300,"candidatesTokenCount":60,"totalTokenCount":360}}]}
integration-tests/browser-agent-localhost.navigate.responses ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll navigate to the localhost page and read its content using the browser agent."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Navigate to http://127.0.0.1:18923/index.html and tell me the page title and list all links on the page"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":40,"totalTokenCount":140}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"navigate_page","args":{"url":"http://127.0.0.1:18923/index.html"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":20,"totalTokenCount":120}}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":150,"candidatesTokenCount":20,"totalTokenCount":170}}]}
4
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"result":{"success":true,"summary":"Page title is 'Test Fixture - Home'. Found 3 links: Contact Form (/form.html), Multi-Step Flow (/multi-step/step1.html), Dynamic Content (/dynamic.html)."}}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":40,"totalTokenCount":240}}]}
5
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"The localhost test fixture page has:\n\n**Title**: Test Fixture - Home\n\n**Links**:\n1. Contact Form (form.html)\n2. Multi-Step Flow (multi-step/step1.html)\n3. Dynamic Content (dynamic.html)\n\nThe page also has a heading 'Test Fixture Home Page' and footer content."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":60,"totalTokenCount":260}}]}
integration-tests/browser-agent.cleanup.responses ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll open https://example.com and check the page title for you."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Open https://example.com and get the page title"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":35,"totalTokenCount":135}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"navigate_page","args":{"url":"https://example.com"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":20,"totalTokenCount":120}}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":150,"candidatesTokenCount":20,"totalTokenCount":170}}]}
4
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"result":{"success":true,"summary":"The page title is 'Example Domain'."}}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":30,"totalTokenCount":230}}]}
5
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I have opened the page and the title is 'Example Domain'. The browser session has been cleaned up successfully."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":30,"totalTokenCount":230}}]}
integration-tests/browser-agent.confirmation.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"file_path":"test.txt","content":"hello"}}},{"text":"I've successfully written \"hello\" to test.txt. The file has been created with the specified content."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
integration-tests/browser-policy.responses ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I'll help you with that."},{"functionCall":{"name":"invoke_agent","args":{"agent_name":"browser_agent","prompt":"Open https://example.com and check if there is a heading"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"new_page","args":{"url":"https://example.com"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"take_snapshot","args":{}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
4
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"complete_task","args":{"success":true,"summary":"SUCCESS_POLICY_TEST_COMPLETED"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":50,"totalTokenCount":150}}]}
5
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Task completed successfully. The page has the heading \"Example Domain\"."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":200,"candidatesTokenCount":50,"totalTokenCount":250}}]}
integration-tests/browser-policy.test.ts ADDED
@@ -0,0 +1,240 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2026 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig, poll } from './test-helper.js';
9
+ import { dirname, join } from 'node:path';
10
+ import { fileURLToPath } from 'node:url';
11
+ import { execSync } from 'node:child_process';
12
+ import { existsSync, writeFileSync, readFileSync, mkdirSync } from 'node:fs';
13
+ import { env } from 'node:process';
14
+ import stripAnsi from 'strip-ansi';
15
+
16
+ // Browser agent Chrome DevTools MCP connection is flaky in Docker sandbox.
17
+ // See: https://github.com/google-gemini/gemini-cli/issues/24382
18
+ const isDockerSandbox = env['GEMINI_SANDBOX'] === 'docker';
19
+
20
+ const __filename = fileURLToPath(import.meta.url);
21
+ const __dirname = dirname(__filename);
22
+
23
+ const chromeAvailable = (() => {
24
+ try {
25
+ if (process.platform === 'darwin') {
26
+ execSync(
27
+ 'test -d "/Applications/Google Chrome.app" || test -d "/Applications/Chromium.app"',
28
+ {
29
+ stdio: 'ignore',
30
+ },
31
+ );
32
+ } else if (process.platform === 'linux') {
33
+ execSync(
34
+ 'which google-chrome || which chromium-browser || which chromium',
35
+ { stdio: 'ignore' },
36
+ );
37
+ } else if (process.platform === 'win32') {
38
+ const chromePaths = [
39
+ 'C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe',
40
+ 'C:\\Program Files (x86)\\Google\\Chrome\\Application\\chrome.exe',
41
+ `${process.env['LOCALAPPDATA'] ?? ''}\\Google\\Chrome\\Application\\chrome.exe`,
42
+ ];
43
+ const found = chromePaths.some((p) => existsSync(p));
44
+ if (!found) {
45
+ execSync('where chrome || where chromium', { stdio: 'ignore' });
46
+ }
47
+ } else {
48
+ return false;
49
+ }
50
+ return true;
51
+ } catch {
52
+ return false;
53
+ }
54
+ })();
55
+
56
+ describe.skipIf(!chromeAvailable)('browser-policy', () => {
57
+ let rig: TestRig;
58
+
59
+ beforeEach(() => {
60
+ rig = new TestRig();
61
+ });
62
+
63
+ afterEach(async () => {
64
+ await rig.cleanup();
65
+ });
66
+
67
+ it.skipIf(isDockerSandbox)(
68
+ 'should skip confirmation when "Allow all server tools for this session" is chosen',
69
+ async () => {
70
+ rig.setup('browser-policy-skip-confirmation', {
71
+ fakeResponsesPath: join(__dirname, 'browser-policy.responses'),
72
+ settings: {
73
+ agents: {
74
+ overrides: {
75
+ browser_agent: {
76
+ enabled: true,
77
+ },
78
+ },
79
+ browser: {
80
+ headless: true,
81
+ sessionMode: 'isolated',
82
+ allowedDomains: ['example.com'],
83
+ },
84
+ },
85
+ },
86
+ });
87
+
88
+ // Manually trust the folder to avoid the dialog and enable option 3
89
+ const geminiDir = join(rig.homeDir!, '.gemini');
90
+ mkdirSync(geminiDir, { recursive: true });
91
+
92
+ // Write to trustedFolders.json
93
+ const trustedFoldersPath = join(geminiDir, 'trustedFolders.json');
94
+ const trustedFolders = {
95
+ [rig.testDir!]: 'TRUST_FOLDER',
96
+ };
97
+ writeFileSync(
98
+ trustedFoldersPath,
99
+ JSON.stringify(trustedFolders, null, 2),
100
+ );
101
+
102
+ // Force confirmation for browser agent.
103
+ // NOTE: We don't force confirm browser tools here because "Allow all server tools"
104
+ // adds a rule with ALWAYS_ALLOW_PRIORITY (3.9x) which would be overshadowed by
105
+ // a rule in the user tier (4.x) like the one from this TOML.
106
+ // By removing the explicit mcp rule, the first MCP tool will still prompt
107
+ // due to default approvalMode = 'default', and then "Allow all" will correctly
108
+ // bypass subsequent tools.
109
+ const policyFile = join(rig.testDir!, 'force-confirm.toml');
110
+ writeFileSync(
111
+ policyFile,
112
+ `
113
+ [[rule]]
114
+ name = "Force confirm browser_agent"
115
+ toolName = "invoke_agent"
116
+ argsPattern = "\\"agent_name\\":\\\\s*\\"browser_agent\\""
117
+ decision = "ask_user"
118
+ priority = 200
119
+ `,
120
+ );
121
+
122
+ // Update settings.json in both project and home directories to point to the policy file
123
+ for (const baseDir of [rig.testDir!, rig.homeDir!]) {
124
+ const settingsPath = join(baseDir, '.gemini', 'settings.json');
125
+ if (existsSync(settingsPath)) {
126
+ const settings = JSON.parse(readFileSync(settingsPath, 'utf-8'));
127
+ settings.policyPaths = [policyFile];
128
+ // Ensure folder trust is enabled
129
+ settings.security = settings.security || {};
130
+ settings.security.folderTrust = settings.security.folderTrust || {};
131
+ settings.security.folderTrust.enabled = true;
132
+ writeFileSync(settingsPath, JSON.stringify(settings, null, 2));
133
+ }
134
+ }
135
+
136
+ const run = await rig.runInteractive({
137
+ approvalMode: 'default',
138
+ env: {
139
+ GEMINI_CLI_INTEGRATION_TEST: 'true',
140
+ },
141
+ });
142
+
143
+ await run.sendKeys(
144
+ 'Open https://example.com and check if there is a heading\r',
145
+ );
146
+ await run.sendKeys('\r');
147
+
148
+ // Handle confirmations.
149
+ // 1. Initial browser_agent delegation (likely only 3 options, so use option 1: Allow once)
150
+ await poll(
151
+ () => stripAnsi(run.output).toLowerCase().includes('action required'),
152
+ 60000,
153
+ 1000,
154
+ );
155
+ await run.sendKeys('1\r');
156
+ await new Promise((r) => setTimeout(r, 2000));
157
+
158
+ // Handle privacy notice
159
+ await poll(
160
+ () => stripAnsi(run.output).toLowerCase().includes('privacy notice'),
161
+ 5000,
162
+ 100,
163
+ );
164
+ await run.sendKeys('1\r');
165
+ await new Promise((r) => setTimeout(r, 5000));
166
+
167
+ // new_page (MCP tool, should have 4 options, use option 3: Allow all server tools)
168
+ await poll(
169
+ () => {
170
+ const stripped = stripAnsi(run.output).toLowerCase();
171
+ return (
172
+ stripped.includes('new_page') &&
173
+ stripped.includes('allow all server tools for this session')
174
+ );
175
+ },
176
+ 60000,
177
+ 1000,
178
+ );
179
+
180
+ // Select "Allow all server tools for this session" (option 3)
181
+ await run.sendKeys('3\r');
182
+
183
+ // Wait for the browser agent to finish (success or failure)
184
+ await poll(
185
+ () => {
186
+ const stripped = stripAnsi(run.output).toLowerCase();
187
+ return (
188
+ stripped.includes('completed successfully') ||
189
+ stripped.includes('agent error')
190
+ );
191
+ },
192
+ 120000,
193
+ 1000,
194
+ );
195
+
196
+ const output = stripAnsi(run.output).toLowerCase();
197
+
198
+ expect(output).toContain('browser_agent');
199
+ // The test validates that "Allow all server tools" skips subsequent
200
+ // tool confirmations — the browser agent may still fail due to
201
+ // Chrome/MCP issues in CI, which is acceptable for this policy test.
202
+ expect(
203
+ output.includes('completed successfully') ||
204
+ output.includes('agent error'),
205
+ ).toBe(true);
206
+ },
207
+ );
208
+
209
+ it('should show the visible warning when browser agent starts in existing session mode', async () => {
210
+ rig.setup('browser-session-warning', {
211
+ fakeResponsesPath: join(__dirname, 'browser-agent.cleanup.responses'),
212
+ settings: {
213
+ general: {
214
+ enableAutoUpdateNotification: false,
215
+ },
216
+ agents: {
217
+ overrides: {
218
+ browser_agent: {
219
+ enabled: true,
220
+ },
221
+ },
222
+ browser: {
223
+ sessionMode: 'existing',
224
+ headless: true,
225
+ },
226
+ },
227
+ },
228
+ });
229
+
230
+ const stdout = await rig.runCommand(['Open https://example.com'], {
231
+ env: {
232
+ GEMINI_API_KEY: 'fake-key',
233
+ GEMINI_TELEMETRY_DISABLED: 'true',
234
+ DEV: 'true',
235
+ },
236
+ });
237
+
238
+ expect(stdout).toContain('saved logins will be visible');
239
+ });
240
+ });
integration-tests/checkpointing.test.ts ADDED
@@ -0,0 +1,155 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import * as fs from 'node:fs/promises';
9
+ import * as path from 'node:path';
10
+ import * as os from 'node:os';
11
+ import { GitService, Storage } from '@google/gemini-cli-core';
12
+
13
+ describe('Checkpointing Integration', () => {
14
+ let tmpDir: string;
15
+ let projectRoot: string;
16
+ let fakeHome: string;
17
+ let originalEnv: NodeJS.ProcessEnv;
18
+
19
+ beforeEach(async () => {
20
+ tmpDir = await fs.mkdtemp(
21
+ path.join(os.tmpdir(), 'gemini-checkpoint-test-'),
22
+ );
23
+ projectRoot = path.join(tmpDir, 'project');
24
+ fakeHome = path.join(tmpDir, 'home');
25
+
26
+ await fs.mkdir(projectRoot, { recursive: true });
27
+ await fs.mkdir(fakeHome, { recursive: true });
28
+
29
+ // Save original env
30
+ originalEnv = { ...process.env };
31
+
32
+ // Simulate environment with NO global gitconfig
33
+ process.env['HOME'] = fakeHome;
34
+ delete process.env['GIT_CONFIG_GLOBAL'];
35
+ delete process.env['GIT_CONFIG_SYSTEM'];
36
+ });
37
+
38
+ afterEach(async () => {
39
+ // Restore env
40
+ process.env = originalEnv;
41
+
42
+ // Cleanup
43
+ try {
44
+ await fs.rm(tmpDir, { recursive: true, force: true });
45
+ } catch (e) {
46
+ console.error('Failed to cleanup temp dir', e);
47
+ }
48
+ });
49
+
50
+ it('should successfully create and restore snapshots without global git config', async () => {
51
+ const storage = new Storage(projectRoot);
52
+ const gitService = new GitService(projectRoot, storage);
53
+
54
+ // 1. Initialize
55
+ await gitService.initialize();
56
+
57
+ // Verify system config empty file creation
58
+ // We need to access getHistoryDir logic or replicate it.
59
+ // Since we don't have access to private getHistoryDir, we can infer it or just trust the functional test.
60
+
61
+ // 2. Create initial state
62
+ await fs.writeFile(path.join(projectRoot, 'file1.txt'), 'version 1');
63
+ await fs.writeFile(path.join(projectRoot, 'file2.txt'), 'permanent file');
64
+
65
+ // 3. Create Snapshot
66
+ const snapshotHash = await gitService.createFileSnapshot('Checkpoint 1');
67
+ expect(snapshotHash).toBeDefined();
68
+
69
+ // 4. Modify files
70
+ await fs.writeFile(
71
+ path.join(projectRoot, 'file1.txt'),
72
+ 'version 2 (BAD CHANGE)',
73
+ );
74
+ await fs.writeFile(
75
+ path.join(projectRoot, 'file3.txt'),
76
+ 'new file (SHOULD BE GONE)',
77
+ );
78
+ await fs.rm(path.join(projectRoot, 'file2.txt'));
79
+
80
+ // 5. Restore
81
+ await gitService.restoreProjectFromSnapshot(snapshotHash);
82
+
83
+ // 6. Verify state
84
+ const file1Content = await fs.readFile(
85
+ path.join(projectRoot, 'file1.txt'),
86
+ 'utf-8',
87
+ );
88
+ expect(file1Content).toBe('version 1');
89
+
90
+ const file2Exists = await fs
91
+ .stat(path.join(projectRoot, 'file2.txt'))
92
+ .then(() => true)
93
+ .catch(() => false);
94
+ expect(file2Exists).toBe(true);
95
+ const file2Content = await fs.readFile(
96
+ path.join(projectRoot, 'file2.txt'),
97
+ 'utf-8',
98
+ );
99
+ expect(file2Content).toBe('permanent file');
100
+
101
+ const file3Exists = await fs
102
+ .stat(path.join(projectRoot, 'file3.txt'))
103
+ .then(() => true)
104
+ .catch(() => false);
105
+ expect(file3Exists).toBe(false);
106
+ });
107
+
108
+ it('should ignore user global git config and use isolated identity', async () => {
109
+ // 1. Create a fake global gitconfig with a specific user
110
+ const globalConfigPath = path.join(fakeHome, '.gitconfig');
111
+ const globalConfigContent = `[user]
112
+ name = Global User
113
+ email = global@example.com
114
+ `;
115
+ await fs.writeFile(globalConfigPath, globalConfigContent);
116
+
117
+ // Point HOME to fakeHome so git picks up this global config (if we didn't isolate it)
118
+ process.env['HOME'] = fakeHome;
119
+ // Ensure GIT_CONFIG_GLOBAL is NOT set for the process initially,
120
+ // so it would default to HOME/.gitconfig if GitService didn't override it.
121
+ delete process.env['GIT_CONFIG_GLOBAL'];
122
+
123
+ const storage = new Storage(projectRoot);
124
+ const gitService = new GitService(projectRoot, storage);
125
+
126
+ await gitService.initialize();
127
+
128
+ // 2. Create a file and snapshot
129
+ await fs.writeFile(path.join(projectRoot, 'test.txt'), 'content');
130
+ await gitService.createFileSnapshot('Snapshot with global config present');
131
+
132
+ // 3. Verify the commit author in the shadow repo
133
+ const historyDir = storage.getHistoryDir();
134
+
135
+ const { execFileSync } = await import('node:child_process');
136
+
137
+ const logOutput = execFileSync(
138
+ 'git',
139
+ ['log', '-1', '--pretty=format:%an <%ae>'],
140
+ {
141
+ cwd: historyDir,
142
+ env: {
143
+ ...process.env,
144
+ GIT_DIR: path.join(historyDir, '.git'),
145
+ GIT_CONFIG_GLOBAL: path.join(historyDir, '.gitconfig'),
146
+ GIT_CONFIG_SYSTEM: path.join(historyDir, '.gitconfig_system_empty'),
147
+ },
148
+ encoding: 'utf-8',
149
+ },
150
+ );
151
+
152
+ expect(logOutput).toBe('Gemini CLI <gemini-cli@google.com>');
153
+ expect(logOutput).not.toContain('Global User');
154
+ });
155
+ });
integration-tests/concurrency-limit.responses ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/1"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/2"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/3"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/4"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/5"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/6"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/7"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/8"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/9"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/10"}}},{"functionCall":{"name":"web_fetch","args":{"prompt":"fetch https://example.com/11"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":100,"candidatesTokenCount":500,"totalTokenCount":600}}]}
2
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 1 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
3
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 2 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
4
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 3 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
5
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 4 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
6
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 5 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
7
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 6 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
8
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 7 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
9
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 8 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
10
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 9 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
11
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Page 10 content"}],"role":"model"},"finishReason":"STOP","index":0}]}}
12
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Some requests were rate limited: Rate limit exceeded for host. Please wait 60 seconds before trying again."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":1000,"candidatesTokenCount":50,"totalTokenCount":1050}}]}
integration-tests/context-compress-interactive.compress-empty.responses ADDED
File without changes
integration-tests/context-compress-interactive.compress-failure.responses ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"thought":true,"text":"**Observing Initial Conditions**\n\nI'm currently focused on the initial context. I've taken note of the provided date, OS, and working directory. I'm also carefully examining the file structure presented within the current working directory. It's helping me understand the starting point for further analysis.\n\n\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12270,"totalTokenCount":12316,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12270}],"thoughtsTokenCount":46}},{"candidates":[{"content":{"parts":[{"thought":true,"text":"**Assessing User Intent**\n\nI'm now shifting my focus. I've successfully registered the provided data and file structure. My current task is to understand the user's ultimate goal, given the information provided. The \"Hello.\" command is straightforward, but I'm checking if there's an underlying objective.\n\n\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12270,"totalTokenCount":12341,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12270}],"thoughtsTokenCount":71}},{"candidates":[{"content":{"parts":[{"thoughtSignature":"CiQB0e2Kb3dRh+BYdbZvmulSN2Pwbc75DfQOT3H4EN0rn039hoMKfwHR7YpvvyqNKoxXAiCbYw3gbcTr/+pegUpgnsIrt8oQPMytFMjKSsMyshfygc21T2MkyuI6Q5I/fNCcHROWexdZnIeppVCDB2TarN4LGW4T9Yci6n/ynMMFT2xc2/vyHpkDgRM7avhMElnBhuxAY+e4TpxkZIncGWCEHP1TouoKpgEB0e2Kb8Xpwm0hiKhPt2ZLizpxjk+CVtcbnlgv69xo5VsuQ+iNyrVGBGRwNx+eTeNGdGpn6e73WOCZeP91FwOZe7URyL12IA6E6gYWqw0kXJR4hO4p6Lwv49E3+FRiG2C4OKDF8LF5XorYyCHSgBFT1/RUAVj81GDTx1xxtmYKN3xq8Ri+HsPbqU/FM/jtNZKkXXAtufw2Bmw8lJfmugENIv/TQI7xCo8BAdHtim8KgAXJfZ7ASfutVLKTylQeaslyB/SmcHJ0ZiNr5j8WP1prZdb6XnZZ1ZNbhjxUf/ymoxHKGvtTPBgLE9azMj8Lx/k0clhd2a+wNsiIqW9qCzlVah0tBMytpQUjIDtQe9Hj4LLUprF9PUe/xJkj000Z0ZzsgFm2ncdTWZTdkhCQDpyETVAxdE+oklwKJAHR7YpvUjSkD6KwY1gLrOsHKy0UNfn2lMbxjVetKNMVBRqsTg==","text":"Hello."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12270,"totalTokenCount":12341,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12270}],"thoughtsTokenCount":71}}]}
2
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"<state_snapshot>\n <overall_goal>\n <!-- The user has not yet specified a goal. -->\n </overall_goal>\n\n <key_knowledge>\n - OS: linux\n - Date: Friday, October 24, 2025\n </key_knowledge>\n\n <file_system_state>\n - OBSERVED: The directory contains `telemetry.log` and a `.gemini/` directory.\n - OBSERVED: The `.gemini/` directory contains `settings.json` and `settings.json.orig`.\n </file_system_state>\n\n <recent_actions>\n - The user initiated the chat.\n </recent_actions>\n\n <current_plan>\n 1. [TODO] Await the user's first instruction to formulate a plan.\n </current_plan>\n</state_snapshot>"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":983,"candidatesTokenCount":299,"totalTokenCount":1637,"promptTokensDetails":[{"modality":"TEXT","tokenCount":983}],"thoughtsTokenCount":355}}}
3
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"<state_snapshot>\n <overall_goal>\n <!-- The user has not yet specified a goal. -->\n </overall_goal>\n\n <key_knowledge>\n - OS: linux\n - Date: Friday, October 24, 2025\n </key_knowledge>\n\n <file_system_state>\n - OBSERVED: The directory contains `telemetry.log` and a `.gemini/` directory.\n - OBSERVED: The `.gemini/` directory contains `settings.json` and `settings.json.orig`.\n </file_system_state>\n\n <recent_actions>\n - The user initiated the chat.\n </recent_actions>\n\n <current_plan>\n 1. [TODO] Await the user's first instruction to formulate a plan.\n </current_plan>\n</state_snapshot>"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":983,"candidatesTokenCount":299,"totalTokenCount":1637,"promptTokensDetails":[{"modality":"TEXT","tokenCount":983}],"thoughtsTokenCount":355}}}
integration-tests/context-fidelity.test.ts ADDED
@@ -0,0 +1,287 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2026 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig } from './test-helper.js';
9
+ import * as path from 'node:path';
10
+ import * as fs from 'node:fs';
11
+ import { FinishReason, GenerateContentResponse } from '@google/genai';
12
+ import type { FakeResponse, HistoryTurn } from '@google/gemini-cli-core';
13
+
14
+ describe('Context Management Fidelity E2E', () => {
15
+ let rig: TestRig;
16
+
17
+ function generateRandomString(length: number): string {
18
+ const characters =
19
+ 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789';
20
+ let result = '';
21
+ for (let i = 0; i < length; i++) {
22
+ result += characters.charAt(
23
+ Math.floor(Math.random() * characters.length),
24
+ );
25
+ }
26
+ return result;
27
+ }
28
+
29
+ beforeEach(() => {
30
+ rig = new TestRig();
31
+ });
32
+
33
+ afterEach(async () => await rig.cleanup());
34
+
35
+ it(
36
+ 'should reproduce the exact context working buffer on resume',
37
+ { timeout: 300000 },
38
+ async () => {
39
+ // Mock responses to trigger GC (summarization)
40
+ const snapshotResponse: FakeResponse = {
41
+ method: 'generateContent',
42
+ response: {
43
+ candidates: [
44
+ {
45
+ content: {
46
+ parts: [
47
+ {
48
+ text: JSON.stringify({
49
+ new_facts: ['GC Triggered.'],
50
+ new_constraints: [],
51
+ new_tasks: [],
52
+ resolved_task_ids: [],
53
+ obsolete_fact_indices: [],
54
+ obsolete_constraint_indices: [],
55
+ chronological_summary: 'Snapshot created.',
56
+ }),
57
+ },
58
+ ],
59
+ role: 'model',
60
+ },
61
+ finishReason: FinishReason.STOP,
62
+ index: 0,
63
+ },
64
+ ],
65
+ } as unknown as GenerateContentResponse,
66
+ };
67
+
68
+ const countTokensResponse: FakeResponse = {
69
+ method: 'countTokens',
70
+ response: { totalTokens: 1000 },
71
+ };
72
+
73
+ const streamResponse = (text: string): FakeResponse => ({
74
+ method: 'generateContentStream',
75
+ response: [
76
+ {
77
+ candidates: [
78
+ {
79
+ content: { parts: [{ text }], role: 'model' },
80
+ finishReason: FinishReason.STOP,
81
+ index: 0,
82
+ },
83
+ ],
84
+ },
85
+ ] as unknown as GenerateContentResponse[],
86
+ });
87
+
88
+ const setupResponses = (fileName: string, mocks: FakeResponse[]) => {
89
+ const filePath = path.join(rig.testDir!, fileName);
90
+ fs.writeFileSync(
91
+ filePath,
92
+ mocks.map((m) => JSON.stringify(m)).join('\n'),
93
+ );
94
+ return filePath;
95
+ };
96
+
97
+ await rig.setup('context-fidelity', {
98
+ settings: {
99
+ experimental: {
100
+ stressTestProfile: true, // Lowers thresholds to trigger GC easily
101
+ },
102
+ },
103
+ });
104
+
105
+ const traceDir = path.join(rig.testDir!, 'traces');
106
+ fs.mkdirSync(traceDir, { recursive: true });
107
+ const traceLog = path.join(traceDir, 'trace.log');
108
+
109
+ // Ignore trace and response files to keep environment context clean and stable
110
+ fs.writeFileSync(
111
+ path.join(rig.testDir!, '.geminiignore'),
112
+ 'traces/\nresp*.json\ndebug.log\n',
113
+ );
114
+
115
+ const commonEnv = {
116
+ GEMINI_API_KEY: 'mock-key',
117
+ GEMINI_CONTEXT_TRACE_DIR: traceDir,
118
+ GEMINI_CONTEXT_TRACE_ENABLED: 'true',
119
+ GEMINI_DEBUG_LOG_FILE: path.join(rig.testDir!, 'debug.log'),
120
+ };
121
+
122
+ const runMocks: FakeResponse[] = [
123
+ streamResponse('Ack 1'),
124
+ streamResponse('Ack 2'),
125
+ streamResponse('Ack 3'),
126
+ streamResponse('Ack 4'),
127
+ streamResponse('Ack 5'),
128
+ streamResponse('Ack 6'),
129
+ streamResponse('Ack 7'),
130
+ streamResponse('Ack 8'),
131
+ streamResponse('Ack 9'),
132
+ streamResponse('Ack 10'),
133
+ streamResponse('Ack 11'),
134
+ streamResponse('Ack 12'),
135
+ ];
136
+ for (let i = 0; i < 50; i++) {
137
+ runMocks.push(snapshotResponse);
138
+ runMocks.push(countTokensResponse);
139
+ }
140
+
141
+ // Turns 1-10: Build up history
142
+ for (let i = 1; i <= 10; i++) {
143
+ await rig.run({
144
+ args: [
145
+ '--debug',
146
+ i === 1 ? '' : '--resume',
147
+ i === 1 ? '' : 'latest',
148
+ '--fake-responses-non-strict',
149
+ setupResponses(`resp_init_${i}.json`, runMocks),
150
+ ].filter(Boolean),
151
+ stdin: `Turn ${i}: ` + generateRandomString(900),
152
+ env: commonEnv,
153
+ });
154
+ }
155
+
156
+ // Turn 11: Penultimate turn
157
+ await rig.run({
158
+ args: [
159
+ '--debug',
160
+ '--resume',
161
+ 'latest',
162
+ '--fake-responses-non-strict',
163
+ setupResponses('resp2.json', runMocks),
164
+ ],
165
+ stdin: 'Turn 11: ' + generateRandomString(900),
166
+ env: commonEnv,
167
+ });
168
+
169
+ // Turn 12: Breach threshold and force GC
170
+ await rig.run({
171
+ args: [
172
+ '--debug',
173
+ '--resume',
174
+ 'latest',
175
+ '--fake-responses-non-strict',
176
+ setupResponses('resp3.json', runMocks),
177
+ ],
178
+ stdin: 'Turn 12: ' + generateRandomString(900),
179
+ env: commonEnv,
180
+ });
181
+
182
+ // Extract the rendered context asset from the log
183
+ const getRenderedContext = (logContent: string): HistoryTurn[] | null => {
184
+ const lines = logContent.split('\n');
185
+ const renderLines = lines.filter(
186
+ (l) =>
187
+ l.includes('[Render] Render Sanitized Context for LLM') ||
188
+ l.includes('[Render] Render Context for LLM'),
189
+ );
190
+ if (renderLines.length === 0) return null;
191
+
192
+ const lastRender = renderLines[renderLines.length - 1];
193
+ const detailsMatch = lastRender.match(/\| Details: (.*)$/);
194
+ if (!detailsMatch) return null;
195
+
196
+ const details = JSON.parse(detailsMatch[1]);
197
+ const assetInfo =
198
+ details.renderedContextSanitized || details.renderedContext;
199
+ if (assetInfo && assetInfo.$asset) {
200
+ const assetPath = path.join(traceDir, 'assets', assetInfo.$asset);
201
+ return JSON.parse(fs.readFileSync(assetPath, 'utf-8'));
202
+ }
203
+ return assetInfo;
204
+ };
205
+
206
+ const log1 = fs.readFileSync(traceLog, 'utf-8');
207
+ const contextBeforeExit = getRenderedContext(log1);
208
+ expect(contextBeforeExit).toBeDefined();
209
+ console.log(
210
+ 'Context Before Exit (First 2 turns):',
211
+ JSON.stringify(contextBeforeExit!.slice(0, 2), null, 2),
212
+ );
213
+
214
+ // Turn 4: Resume and run a small command
215
+ await rig.run({
216
+ args: [
217
+ '--debug',
218
+ '--resume',
219
+ 'latest',
220
+ '--fake-responses-non-strict',
221
+ setupResponses('resp4.json', runMocks),
222
+ 'continue',
223
+ ],
224
+ env: commonEnv,
225
+ });
226
+
227
+ const log2 = fs.readFileSync(traceLog, 'utf-8');
228
+ const contextAfterResume = getRenderedContext(log2);
229
+ expect(contextAfterResume).toBeDefined();
230
+ console.log(
231
+ 'Context After Resume (First 2 turns):',
232
+ JSON.stringify(contextAfterResume!.slice(0, 2), null, 2),
233
+ );
234
+
235
+ expect(contextAfterResume!.length).toBeGreaterThanOrEqual(
236
+ contextBeforeExit!.length,
237
+ );
238
+
239
+ // The environment context is intentionally refreshed on resume to reflect
240
+ // the current state of the workspace (e.g. new files, current date).
241
+ // We allow its content to differ but ensure it's still an environment context.
242
+ const isEnvContext = (turn: HistoryTurn) =>
243
+ turn.content.parts?.some((p) => p.text?.includes('<session_context>'));
244
+
245
+ for (let i = 0; i < contextBeforeExit!.length; i++) {
246
+ expect(contextAfterResume![i].id).toBe(contextBeforeExit![i].id);
247
+
248
+ const turnBefore = contextBeforeExit![i];
249
+ const turnAfter = contextAfterResume![i];
250
+
251
+ if (isEnvContext(turnBefore)) {
252
+ expect(isEnvContext(turnAfter)).toBe(true);
253
+ continue;
254
+ }
255
+
256
+ expect(turnAfter.content).toEqual(turnBefore.content);
257
+ }
258
+
259
+ // Most importantly, synthetic IDs (like summaries) must be stable.
260
+ const syntheticTurns = contextBeforeExit!.filter(
261
+ (t: HistoryTurn) =>
262
+ t.content.parts?.some((p) => p.text?.includes('active_tasks')) ||
263
+ (t.id && t.id.length === 32),
264
+ );
265
+ expect(syntheticTurns.length).toBeGreaterThan(0);
266
+
267
+ const syntheticTurnsAfter = contextAfterResume!.filter(
268
+ (t: HistoryTurn) =>
269
+ t.content.parts?.some((p) => p.text?.includes('active_tasks')) ||
270
+ (t.id && t.id.length === 32),
271
+ );
272
+ expect(syntheticTurnsAfter.length).toBeGreaterThanOrEqual(
273
+ syntheticTurns.length,
274
+ );
275
+
276
+ // Check if the first synthetic turn is identical (with relaxation for environment context)
277
+ expect(syntheticTurnsAfter[0].id).toBe(syntheticTurns[0].id);
278
+ if (isEnvContext(syntheticTurns[0])) {
279
+ expect(isEnvContext(syntheticTurnsAfter[0])).toBe(true);
280
+ } else {
281
+ expect(syntheticTurnsAfter[0].content).toEqual(
282
+ syntheticTurns[0].content,
283
+ );
284
+ }
285
+ },
286
+ );
287
+ });
integration-tests/extensions-install.test.ts ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, expect, it, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig } from './test-helper.js';
9
+ import { writeFileSync } from 'node:fs';
10
+ import { join } from 'node:path';
11
+
12
+ const extension = `{
13
+ "name": "test-extension-install",
14
+ "version": "0.0.1"
15
+ }`;
16
+
17
+ const extensionUpdate = `{
18
+ "name": "test-extension-install",
19
+ "version": "0.0.2"
20
+ }`;
21
+
22
+ describe('extension install', () => {
23
+ let rig: TestRig;
24
+
25
+ beforeEach(() => {
26
+ rig = new TestRig();
27
+ });
28
+
29
+ afterEach(async () => await rig.cleanup());
30
+
31
+ it('installs a local extension, verifies a command, and updates it', async () => {
32
+ rig.setup('extension install test');
33
+ const testServerPath = join(rig.testDir!, 'gemini-extension.json');
34
+ writeFileSync(testServerPath, extension);
35
+ try {
36
+ const result = await rig.runCommand(
37
+ ['--debug', 'extensions', 'install', `${rig.testDir!}`],
38
+ { stdin: 'y\n' },
39
+ );
40
+ expect(result).toContain('test-extension-install');
41
+
42
+ const listResult = await rig.runCommand([
43
+ '--debug',
44
+ 'extensions',
45
+ 'list',
46
+ ]);
47
+ expect(listResult).toContain('test-extension-install');
48
+ writeFileSync(testServerPath, extensionUpdate);
49
+ const updateResult = await rig.runCommand(
50
+ ['--debug', 'extensions', 'update', `test-extension-install`],
51
+ { stdin: 'y\n' },
52
+ );
53
+ expect(updateResult).toContain('0.0.2');
54
+ } finally {
55
+ await rig.runCommand([
56
+ 'extensions',
57
+ 'uninstall',
58
+ 'test-extension-install',
59
+ ]);
60
+ }
61
+ });
62
+ });
integration-tests/extensions-reload.test.ts ADDED
@@ -0,0 +1,151 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { expect, it, describe, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig } from './test-helper.js';
9
+ import { TestMcpServer } from './test-mcp-server.js';
10
+ import { writeFileSync } from 'node:fs';
11
+ import { join } from 'node:path';
12
+ import { safeJsonStringify } from '@google/gemini-cli-core/src/utils/safeJsonStringify.js';
13
+
14
+ import stripAnsi from 'strip-ansi';
15
+
16
+ describe('extension reloading', () => {
17
+ let rig: TestRig;
18
+
19
+ beforeEach(() => {
20
+ rig = new TestRig();
21
+ });
22
+
23
+ afterEach(async () => await rig.cleanup());
24
+
25
+ // always fails
26
+ // TODO(#14527): Re-enable this once fixed
27
+ it.skip('installs a local extension, updates it, checks it was reloaded properly', async () => {
28
+ const serverA = new TestMcpServer();
29
+ const portA = await serverA.start({
30
+ hello: () => ({ content: [{ type: 'text', text: 'world' }] }),
31
+ });
32
+ const extension = {
33
+ name: 'test-extension',
34
+ version: '0.0.1',
35
+ mcpServers: {
36
+ 'test-server': {
37
+ httpUrl: `http://localhost:${portA}/mcp`,
38
+ },
39
+ },
40
+ };
41
+
42
+ rig.setup('extension reload test', {
43
+ settings: {
44
+ experimental: { extensionReloading: true },
45
+ },
46
+ });
47
+ const testServerPath = join(rig.testDir!, 'gemini-extension.json');
48
+ writeFileSync(testServerPath, safeJsonStringify(extension, 2));
49
+ // defensive cleanup from previous tests.
50
+ try {
51
+ await rig.runCommand(['extensions', 'uninstall', 'test-extension']);
52
+ } catch {
53
+ /* empty */
54
+ }
55
+
56
+ const result = await rig.runCommand(
57
+ ['--debug', 'extensions', 'install', `${rig.testDir!}`],
58
+ { stdin: 'y\n' },
59
+ );
60
+ expect(result).toContain('test-extension');
61
+
62
+ // Now create the update, but its not installed yet
63
+ const serverB = new TestMcpServer();
64
+ const portB = await serverB.start({
65
+ goodbye: () => ({ content: [{ type: 'text', text: 'world' }] }),
66
+ });
67
+ extension.version = '0.0.2';
68
+ extension.mcpServers['test-server'].httpUrl =
69
+ `http://localhost:${portB}/mcp`;
70
+ writeFileSync(testServerPath, safeJsonStringify(extension, 2));
71
+
72
+ // Start the CLI.
73
+ const run = await rig.runInteractive({ args: '--debug' });
74
+ await run.expectText('You have 1 extension with an update available');
75
+ // See the outdated extension
76
+ await run.sendText('/extensions list');
77
+ await run.type('\r');
78
+ await run.expectText('test-extension (v0.0.1) - active (update available)');
79
+ // Wait for the UI to settle and retry the command until we see the update
80
+ await new Promise((resolve) => setTimeout(resolve, 1000));
81
+
82
+ // Poll for the updated list
83
+ await rig.pollCommand(
84
+ async () => {
85
+ await run.sendText('/mcp list');
86
+ await run.type('\r');
87
+ },
88
+ () => {
89
+ const output = stripAnsi(run.output);
90
+ return (
91
+ output.includes(
92
+ 'test-server (from test-extension) - Ready (1 tool)',
93
+ ) && output.includes('- mcp_test-server_hello')
94
+ );
95
+ },
96
+ 30000, // 30s timeout
97
+ );
98
+
99
+ // Update the extension, expect the list to update, and mcp servers as well.
100
+ await run.sendKeys('\u0015/extensions update test-extension');
101
+ await run.expectText('/extensions update test-extension');
102
+ await run.type('\r');
103
+ await new Promise((resolve) => setTimeout(resolve, 500));
104
+ await run.type('\r');
105
+ await run.expectText(
106
+ ` * test-server (remote): http://localhost:${portB}/mcp`,
107
+ );
108
+ await run.type('\r'); // consent
109
+ await run.expectText(
110
+ 'Extension "test-extension" successfully updated: 0.0.1 → 0.0.2',
111
+ );
112
+
113
+ // Poll for the updated extension version
114
+ await rig.pollCommand(
115
+ async () => {
116
+ await run.sendText('/extensions list');
117
+ await run.type('\r');
118
+ },
119
+ () =>
120
+ stripAnsi(run.output).includes(
121
+ 'test-extension (v0.0.2) - active (updated)',
122
+ ),
123
+ 30000,
124
+ );
125
+
126
+ // Poll for the updated mcp tool
127
+ await rig.pollCommand(
128
+ async () => {
129
+ await run.sendText('/mcp list');
130
+ await run.type('\r');
131
+ },
132
+ () => {
133
+ const output = stripAnsi(run.output);
134
+ return (
135
+ output.includes(
136
+ 'test-server (from test-extension) - Ready (1 tool)',
137
+ ) && output.includes('- mcp_test-server_goodbye')
138
+ );
139
+ },
140
+ 30000,
141
+ );
142
+
143
+ await run.sendText('/quit');
144
+ await run.type('\r');
145
+
146
+ // Clean things up.
147
+ await serverA.stop();
148
+ await serverB.stop();
149
+ await rig.runCommand(['extensions', 'uninstall', 'test-extension']);
150
+ });
151
+ });
integration-tests/file-system-interactive.test.ts ADDED
@@ -0,0 +1,67 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { expect, describe, it, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig, skipFlaky } from './test-helper.js';
9
+
10
+ describe.skipIf(skipFlaky)('Interactive file system', () => {
11
+ let rig: TestRig;
12
+
13
+ beforeEach(() => {
14
+ rig = new TestRig();
15
+ });
16
+
17
+ afterEach(async () => {
18
+ await rig.cleanup();
19
+ });
20
+
21
+ it('should perform a read-then-write sequence', async () => {
22
+ const fileName = 'version.txt';
23
+ await rig.setup('interactive-read-then-write', {
24
+ settings: {
25
+ security: {
26
+ auth: {
27
+ selectedType: 'gemini-api-key',
28
+ },
29
+ disableYoloMode: false,
30
+ },
31
+ },
32
+ });
33
+ rig.createFile(fileName, '1.0.0');
34
+
35
+ const run = await rig.runInteractive({
36
+ env: {
37
+ GEMINI_CLI_TRUST_WORKSPACE: 'true',
38
+ },
39
+ });
40
+
41
+ // Step 1: Read the file
42
+ const readPrompt = `Read the version from ${fileName} using the read_file tool`;
43
+ await run.type(readPrompt);
44
+ await run.type('\r');
45
+
46
+ const readCall = await rig.waitForToolCall('read_file', 30000);
47
+ expect(readCall, 'Expected to find a read_file tool call').toBe(true);
48
+
49
+ // Wait for the CLI to finish outputting the response and show the prompt again
50
+ await run.expectText('Type your message', 30000);
51
+
52
+ // Step 2: Write the file
53
+ const writePrompt = `now change the version to 1.0.1 in ${fileName} using the write_file tool`;
54
+ await run.type(writePrompt);
55
+ await run.type('\r');
56
+
57
+ // Check tool calls made with right args
58
+ await rig.expectToolCallSuccess(
59
+ ['write_file', 'replace'],
60
+ 30000,
61
+ (args) => args.includes('1.0.1') && args.includes(fileName),
62
+ );
63
+
64
+ // Wait for telemetry to flush and file system to sync, especially in sandboxed environments
65
+ await rig.waitForTelemetryReady();
66
+ }, 120000);
67
+ });
integration-tests/flicker-detector.max-height.responses ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"{\n \"reasoning\": \"The user is asking for a simple piece of information ('a fun fact'). This is a direct, bounded request with low operational complexity and does not require strategic planning, extensive investigation, or debugging.\",\n \"model_choice\": \"flash\"\n}"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":1173,"candidatesTokenCount":59,"totalTokenCount":1344,"promptTokensDetails":[{"modality":"TEXT","tokenCount":1173}],"thoughtsTokenCount":112}}}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"thought":true,"text":"**Locating a fun fact**\n\nI'm now searching for a fun fact using the web search tool, focusing on finding something engaging and potentially surprising. The goal is to provide a brief, interesting piece of information.\n\n\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12226,"totalTokenCount":12255,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12226}],"thoughtsTokenCount":29}},{"candidates":[{"content":{"parts":[{"thoughtSignature":"CikB0e2Kb1vYSbIdmBfclWY7z4mOZgPxUGi3CtNXYYV9CSmG+SpVXZZkmQpZAdHtim9HVruyrUZZcHKDvIfn3j6/zLMgepC4Pqd79pG641PkPJnnCqEfVFRxmE2NX3Tj2lwRhtuIYT9Cc3CfvWGjbuuvwzynMCApxpIvxdXac/fXJYeRHTsKQQHR7Ypv6eOvWUFUTRGm1x29v8ZnGjtudG31H/Dgc65Y47c594ZJfX9RqJJil0I52Bxsm8UQ74rbARqwT7zYEbNO","functionCall":{"name":"google_web_search","args":{"query":"fun fact"}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12226,"candidatesTokenCount":17,"totalTokenCount":12272,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12226}],"thoughtsTokenCount":29}}]}
3
+ {"method":"generateContent","response":{"candidates":[{"content":{"parts":[{"text":"Here's a fun fact: A day on Venus is longer than a year on Venus. It takes approximately 243 Earth days for Venus to rotate once on its axis, while its orbit around the Sun is about 225 Earth days."}],"role":"model"},"finishReason":"STOP","groundingMetadata":{"searchEntryPoint":{"renderedContent":"<style>\n.container {\n align-items: center;\n border-radius: 8px;\n display: flex;\n font-family: Google Sans, Roboto, sans-serif;\n font-size: 14px;\n line-height: 20px;\n padding: 8px 12px;\n}\n.chip {\n display: inline-block;\n border: solid 1px;\n border-radius: 16px;\n min-width: 14px;\n padding: 5px 16px;\n text-align: center;\n user-select: none;\n margin: 0 8px;\n -webkit-tap-highlight-color: transparent;\n}\n.carousel {\n overflow: auto;\n scrollbar-width: none;\n white-space: nowrap;\n margin-right: -12px;\n}\n.headline {\n display: flex;\n margin-right: 4px;\n}\n.gradient-container {\n position: relative;\n}\n.gradient {\n position: absolute;\n transform: translate(3px, -9px);\n height: 36px;\n width: 9px;\n}\n@media (prefers-color-scheme: light) {\n .container {\n background-color: #fafafa;\n box-shadow: 0 0 0 1px #0000000f;\n }\n .headline-label {\n color: #1f1f1f;\n }\n .chip {\n background-color: #ffffff;\n border-color: #d2d2d2;\n color: #5e5e5e;\n text-decoration: none;\n }\n .chip:hover {\n background-color: #f2f2f2;\n }\n .chip:focus {\n background-color: #f2f2f2;\n }\n .chip:active {\n background-color: #d8d8d8;\n border-color: #b6b6b6;\n }\n .logo-dark {\n display: none;\n }\n .gradient {\n background: linear-gradient(90deg, #fafafa 15%, #fafafa00 100%);\n }\n}\n@media (prefers-color-scheme: dark) {\n .container {\n background-color: #1f1f1f;\n box-shadow: 0 0 0 1px #ffffff26;\n }\n .headline-label {\n color: #fff;\n }\n .chip {\n background-color: #2c2c2c;\n border-color: #3c4043;\n color: #fff;\n text-decoration: none;\n }\n .chip:hover {\n background-color: #353536;\n }\n .chip:focus {\n background-color: #353536;\n }\n .chip:active {\n background-color: #464849;\n border-color: #53575b;\n }\n .logo-light {\n display: none;\n }\n .gradient {\n background: linear-gradient(90deg, #1f1f1f 15%, #1f1f1f00 100%);\n }\n}\n</style>\n<div class=\"container\">\n <div class=\"headline\">\n <svg class=\"logo-light\" width=\"18\" height=\"18\" viewBox=\"9 9 35 35\" fill=\"none\" xmlns=\"http://www.w3.org/2000/svg\">\n <path fill-rule=\"evenodd\" clip-rule=\"evenodd\" d=\"M42.8622 27.0064C42.8622 25.7839 42.7525 24.6084 42.5487 23.4799H26.3109V30.1568H35.5897C35.1821 32.3041 33.9596 34.1222 32.1258 35.3448V39.6864H37.7213C40.9814 36.677 42.8622 32.2571 42.8622 27.0064V27.0064Z\" fill=\"#4285F4\"/>\n <path fill-rule=\"evenodd\" clip-rule=\"evenodd\" d=\"M26.3109 43.8555C30.9659 43.8555 34.8687 42.3195 37.7213 39.6863L32.1258 35.3447C30.5898 36.3792 28.6306 37.0061 26.3109 37.0061C21.8282 37.0061 18.0195 33.9811 16.6559 29.906H10.9194V34.3573C13.7563 39.9841 19.5712 43.8555 26.3109 43.8555V43.8555Z\" fill=\"#34A853\"/>\n <path fill-rule=\"evenodd\" clip-rule=\"evenodd\" d=\"M16.6559 29.8904C16.3111 28.8559 16.1074 27.7588 16.1074 26.6146C16.1074 25.4704 16.3111 24.3733 16.6559 23.3388V18.8875H10.9194C9.74388 21.2072 9.06992 23.8247 9.06992 26.6146C9.06992 29.4045 9.74388 32.022 10.9194 34.3417L15.3864 30.8621L16.6559 29.8904V29.8904Z\" fill=\"#FBBC05\"/>\n <path fill-rule=\"evenodd\" clip-rule=\"evenodd\" d=\"M26.3109 16.2386C28.85 16.2386 31.107 17.1164 32.9095 18.8091L37.8466 13.8719C34.853 11.082 30.9659 9.3736 26.3109 9.3736C19.5712 9.3736 13.7563 13.245 10.9194 18.8875L16.6559 23.3388C18.0195 19.2636 21.8282 16.2386 26.3109 16.2386V16.2386Z\" fill=\"#EA4335\"/>\n </svg>\n <svg class=\"logo-dark\" width=\"18\" height=\"18\" viewBox=\"0 0 48 48\" xmlns=\"http://www.w3.org/2000/svg\">\n <circle cx=\"24\" cy=\"23\" fill=\"#FFF\" r=\"22\"/>\n <path d=\"M33.76 34.26c2.75-2.56 4.49-6.37 4.49-11.26 0-.89-.08-1.84-.29-3H24.01v5.99h8.03c-.4 2.02-1.5 3.56-3.07 4.56v.75l3.91 2.97h.88z\" fill=\"#4285F4\"/>\n <path d=\"M15.58 25.77A8.845 8.845 0 0 0 24 31.86c1.92 0 3.62-.46 4.97-1.31l4.79 3.71C31.14 36.7 27.65 38 24 38c-5.93 0-11.01-3.4-13.45-8.36l.17-1.01 4.06-2.85h.8z\" fill=\"#34A853\"/>\n <path d=\"M15.59 20.21a8.864 8.864 0 0 0 0 5.58l-5.03 3.86c-.98-2-1.53-4.25-1.53-6.64 0-2.39.55-4.64 1.53-6.64l1-.22 3.81 2.98.22 1.08z\" fill=\"#FBBC05\"/>\n <path d=\"M24 14.14c2.11 0 4.02.75 5.52 1.98l4.36-4.36C31.22 9.43 27.81 8 24 8c-5.93 0-11.01 3.4-13.45 8.36l5.03 3.85A8.86 8.86 0 0 1 24 14.14z\" fill=\"#EA4335\"/>\n </svg>\n <div class=\"gradient-container\"><div class=\"gradient\"></div></div>\n </div>\n <div class=\"carousel\">\n <a class=\"chip\" href=\"https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGn08vWlJZ4WluVTi_HlxdQkeXfoN9NWENM8cEINX-BCIIAjcsUGPJ6fPYpoDZM8jiOnbW3cfNip1ONou6w0w34KxnYlV8uNgO8fzTZTkxAcORxmy0KeaUnVbKd6AL6i8M05TqIWCzB4flc3XIEtwVAYStd5HFtahr75GNSZ_VzV1mD1POLYD2rwTfT\">fun fact</a>\n </div>\n</div>\n"},"groundingChunks":[{"web":{"uri":"https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQF-NBVWZeEqhT2BixBuiSaCHeF50iewha2f3M2FfpiNsStuPxhc3sLEzXLR7IFsBbzUBO2kbUmm-usnToWabMSvOIT4ZnTXedj5ZkpwFlYyuadyuBhLNKKJQtOGgg9JTNiwvKxBWt2beHYUjelTJXfVPb0Iy8SVJTahtA3GDA==","title":"hellosubs.co"}}],"groundingSupports":[{"segment":{"startIndex":66,"endIndex":197,"text":"It takes approximately 243 Earth days for Venus to rotate once on its axis, while its orbit around the Sun is about 225 Earth days."},"groundingChunkIndices":[0]}],"webSearchQueries":["fun fact"]},"index":0}],"usageMetadata":{"promptTokenCount":8186,"candidatesTokenCount":65,"totalTokenCount":16468,"cachedContentTokenCount":5360,"promptTokensDetails":[{"modality":"TEXT","tokenCount":8186}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":5360}],"toolUsePromptTokenCount":8207,"toolUsePromptTokensDetails":[{"modality":"TEXT","tokenCount":8207}],"thoughtsTokenCount":10}}}
integration-tests/globalSetup.ts ADDED
@@ -0,0 +1,155 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ // Unset NO_COLOR environment variable to ensure consistent theme behavior between local and CI test runs
8
+ if (process.env['NO_COLOR'] !== undefined) {
9
+ delete process.env['NO_COLOR'];
10
+ }
11
+
12
+ import { mkdir, readdir, rm, readFile } from 'node:fs/promises';
13
+ import { join, dirname, extname } from 'node:path';
14
+ import { fileURLToPath } from 'node:url';
15
+ import { resolveRipgrepPath } from '../packages/core/src/tools/ripGrep.js';
16
+ import { disableMouseTracking } from '@google/gemini-cli-core';
17
+ import { isolateTestEnv } from '../packages/test-utils/src/env-setup.js';
18
+ import { createServer, type Server } from 'node:http';
19
+
20
+ const __dirname = dirname(fileURLToPath(import.meta.url));
21
+ const rootDir = join(__dirname, '..');
22
+ const integrationTestsDir = join(rootDir, '.integration-tests');
23
+ let runDir = ''; // Make runDir accessible in teardown
24
+ let fixtureServer: Server | undefined;
25
+
26
+ const FIXTURE_PORT = 18923;
27
+ const FIXTURE_DIR = join(__dirname, 'test-fixtures');
28
+
29
+ const MIME_TYPES: Record<string, string> = {
30
+ '.html': 'text/html',
31
+ '.css': 'text/css',
32
+ '.js': 'application/javascript',
33
+ '.json': 'application/json',
34
+ '.png': 'image/png',
35
+ '.jpg': 'image/jpeg',
36
+ '.svg': 'image/svg+xml',
37
+ };
38
+
39
+ async function startFixtureServer(): Promise<number> {
40
+ return new Promise((resolve, reject) => {
41
+ const server = createServer(async (req, res) => {
42
+ const urlPath = req.url?.split('?')[0] || '/';
43
+ const relativePath = urlPath === '/' ? 'index.html' : urlPath;
44
+ const filePath = join(FIXTURE_DIR, relativePath);
45
+
46
+ if (!filePath.startsWith(FIXTURE_DIR)) {
47
+ res.writeHead(403, { 'Content-Type': 'text/html' });
48
+ res.end('<h1>403 Forbidden</h1>');
49
+ return;
50
+ }
51
+
52
+ try {
53
+ const content = await readFile(filePath);
54
+ const ext = extname(filePath);
55
+ res.writeHead(200, {
56
+ 'Content-Type': MIME_TYPES[ext] || 'application/octet-stream',
57
+ });
58
+ res.end(content);
59
+ } catch {
60
+ res.writeHead(404, { 'Content-Type': 'text/html' });
61
+ res.end('<h1>404 Not Found</h1>');
62
+ }
63
+ });
64
+
65
+ server.on('error', (err: NodeJS.ErrnoException) => {
66
+ if (err.code === 'EADDRINUSE') {
67
+ console.warn(
68
+ `Port ${FIXTURE_PORT} in use, trying ${FIXTURE_PORT + 1}...`,
69
+ );
70
+ server.listen(FIXTURE_PORT + 1, '127.0.0.1');
71
+ } else {
72
+ reject(err);
73
+ }
74
+ });
75
+
76
+ server.on('listening', () => {
77
+ const addr = server.address();
78
+ const port = typeof addr === 'object' && addr ? addr.port : FIXTURE_PORT;
79
+ fixtureServer = server;
80
+ console.log(`Test fixture server listening on http://127.0.0.1:${port}`);
81
+ resolve(port);
82
+ });
83
+
84
+ server.listen(FIXTURE_PORT, '127.0.0.1');
85
+ });
86
+ }
87
+
88
+ export async function setup() {
89
+ runDir = join(integrationTestsDir, `${Date.now()}`);
90
+ await mkdir(runDir, { recursive: true });
91
+
92
+ // Isolate environment variables
93
+ isolateTestEnv(runDir);
94
+
95
+ // Download ripgrep to avoid race conditions in parallel tests
96
+ const available = await resolveRipgrepPath();
97
+ if (!available) {
98
+ throw new Error('Failed to download ripgrep binary');
99
+ }
100
+
101
+ // Start the test fixture server
102
+ const port = await startFixtureServer();
103
+ process.env['TEST_FIXTURE_PORT'] = String(port);
104
+
105
+ // Clean up old test runs, but keep the latest few for debugging
106
+ try {
107
+ const testRuns = await readdir(integrationTestsDir);
108
+ if (testRuns.length > 5) {
109
+ const oldRuns = testRuns.sort().slice(0, testRuns.length - 5);
110
+ await Promise.all(
111
+ oldRuns.map((oldRun) =>
112
+ rm(join(integrationTestsDir, oldRun), {
113
+ recursive: true,
114
+ force: true,
115
+ }),
116
+ ),
117
+ );
118
+ }
119
+ } catch (e) {
120
+ console.error('Error cleaning up old test runs:', e);
121
+ }
122
+
123
+ process.env['INTEGRATION_TEST_FILE_DIR'] = runDir;
124
+
125
+ if (process.env['KEEP_OUTPUT']) {
126
+ console.log(`Keeping output for test run in: ${runDir}`);
127
+ }
128
+ process.env['VERBOSE'] = process.env['VERBOSE'] ?? 'false';
129
+
130
+ console.log(`\nIntegration test output directory: ${runDir}`);
131
+ }
132
+
133
+ export async function teardown() {
134
+ // Stop the fixture server
135
+ if (fixtureServer) {
136
+ await new Promise<void>((resolve) => {
137
+ fixtureServer!.close(() => resolve());
138
+ });
139
+ fixtureServer = undefined;
140
+ }
141
+
142
+ // Disable mouse tracking
143
+ if (process.stdout.isTTY) {
144
+ disableMouseTracking();
145
+ }
146
+
147
+ // Cleanup the test run directory unless KEEP_OUTPUT is set
148
+ if (process.env['KEEP_OUTPUT'] !== 'true' && runDir) {
149
+ try {
150
+ await rm(runDir, { recursive: true, force: true });
151
+ } catch (e) {
152
+ console.warn('Failed to clean up test run directory:', e);
153
+ }
154
+ }
155
+ }
integration-tests/google_web_search.test.ts ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { WEB_SEARCH_TOOL_NAME } from '../packages/core/src/tools/tool-names.js';
8
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
9
+ import {
10
+ TestRig,
11
+ printDebugInfo,
12
+ assertModelHasOutput,
13
+ checkModelOutputContent,
14
+ } from './test-helper.js';
15
+
16
+ describe('web search tool', () => {
17
+ let rig: TestRig;
18
+
19
+ beforeEach(() => {
20
+ rig = new TestRig();
21
+ });
22
+
23
+ afterEach(async () => await rig.cleanup());
24
+
25
+ it('should be able to search the web', async () => {
26
+ await rig.setup('should be able to search the web', {
27
+ settings: { tools: { core: [WEB_SEARCH_TOOL_NAME] } },
28
+ });
29
+
30
+ let result;
31
+ try {
32
+ result = await rig.run({ args: `what is the weather in London` });
33
+ } catch (error) {
34
+ // Network errors can occur in CI environments
35
+ if (
36
+ error instanceof Error &&
37
+ (error.message.includes('network') || error.message.includes('timeout'))
38
+ ) {
39
+ console.warn(
40
+ 'Skipping test due to network error:',
41
+ (error as Error).message,
42
+ );
43
+ return; // Skip the test
44
+ }
45
+ throw error; // Re-throw if not a network error
46
+ }
47
+
48
+ const foundToolCall = await rig.waitForToolCall(WEB_SEARCH_TOOL_NAME);
49
+
50
+ // Add debugging information
51
+ if (!foundToolCall) {
52
+ const allTools = printDebugInfo(rig, result);
53
+
54
+ // Check if the tool call failed due to network issues
55
+ const failedSearchCalls = allTools.filter(
56
+ (t) =>
57
+ t.toolRequest.name === WEB_SEARCH_TOOL_NAME && !t.toolRequest.success,
58
+ );
59
+ if (failedSearchCalls.length > 0) {
60
+ console.warn(
61
+ `${WEB_SEARCH_TOOL_NAME} tool was called but failed, possibly due to network issues`,
62
+ );
63
+ console.warn(
64
+ 'Failed calls:',
65
+ failedSearchCalls.map((t) => t.toolRequest.args),
66
+ );
67
+ return; // Skip the test if network issues
68
+ }
69
+ }
70
+
71
+ expect(
72
+ foundToolCall,
73
+ `Expected to find a call to ${WEB_SEARCH_TOOL_NAME}`,
74
+ ).toBeTruthy();
75
+
76
+ assertModelHasOutput(result);
77
+ const hasExpectedContent = checkModelOutputContent(result, {
78
+ expectedContent: ['weather', 'london'],
79
+ testName: 'Google web search test',
80
+ });
81
+
82
+ // If content was missing, log the search queries used
83
+ if (!hasExpectedContent) {
84
+ const searchCalls = rig
85
+ .readToolLogs()
86
+ .filter((t) => t.toolRequest.name === WEB_SEARCH_TOOL_NAME);
87
+ if (searchCalls.length > 0) {
88
+ console.warn(
89
+ 'Search queries used:',
90
+ searchCalls.map((t) => t.toolRequest.args),
91
+ );
92
+ }
93
+ }
94
+ });
95
+ });
integration-tests/hooks-agent-flow-multistep.responses ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"list_dir","args":{"path":"."}}}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":10,"candidatesTokenCount":10,"totalTokenCount":20}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Final Answer"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":10,"candidatesTokenCount":10,"totalTokenCount":20}}]}
integration-tests/hooks-agent-flow.test.ts ADDED
@@ -0,0 +1,338 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * @license
3
+ * Copyright 2025 Google LLC
4
+ * SPDX-License-Identifier: Apache-2.0
5
+ */
6
+
7
+ import { describe, it, expect, beforeEach, afterEach } from 'vitest';
8
+ import { TestRig, normalizePath } from './test-helper.js';
9
+ import { join } from 'node:path';
10
+ import { writeFileSync } from 'node:fs';
11
+
12
+ describe('Hooks Agent Flow', () => {
13
+ let rig: TestRig;
14
+
15
+ beforeEach(() => {
16
+ rig = new TestRig();
17
+ });
18
+
19
+ afterEach(async () => {
20
+ if (rig) {
21
+ await rig.cleanup();
22
+ }
23
+ });
24
+
25
+ describe('BeforeAgent Hooks', () => {
26
+ it('should inject additional context via BeforeAgent hook', async () => {
27
+ await rig.setup('should inject additional context via BeforeAgent hook', {
28
+ fakeResponsesPath: join(
29
+ import.meta.dirname,
30
+ 'hooks-agent-flow.responses',
31
+ ),
32
+ });
33
+
34
+ const hookScript = `
35
+ try {
36
+ const output = {
37
+ decision: "allow",
38
+ hookSpecificOutput: {
39
+ hookEventName: "BeforeAgent",
40
+ additionalContext: "SYSTEM INSTRUCTION: This is injected context."
41
+ }
42
+ };
43
+ process.stdout.write(JSON.stringify(output));
44
+ } catch (e) {
45
+ console.error('Failed to write stdout:', e);
46
+ process.exit(1);
47
+ }
48
+ console.error('DEBUG: BeforeAgent hook executed');
49
+ `;
50
+
51
+ const scriptPath = join(rig.testDir!, 'before_agent_context.cjs');
52
+ writeFileSync(scriptPath, hookScript);
53
+
54
+ await rig.setup('should inject additional context via BeforeAgent hook', {
55
+ settings: {
56
+ hooksConfig: {
57
+ enabled: true,
58
+ },
59
+ hooks: {
60
+ BeforeAgent: [
61
+ {
62
+ hooks: [
63
+ {
64
+ type: 'command',
65
+ command: `node "${scriptPath}"`,
66
+ timeout: 5000,
67
+ },
68
+ ],
69
+ },
70
+ ],
71
+ },
72
+ },
73
+ });
74
+
75
+ await rig.run({ args: 'Hello test' });
76
+
77
+ // Verify hook execution and telemetry
78
+ const hookTelemetryFound = await rig.waitForTelemetryEvent('hook_call');
79
+ expect(hookTelemetryFound).toBeTruthy();
80
+
81
+ const hookLogs = rig.readHookLogs();
82
+ const beforeAgentLog = hookLogs.find(
83
+ (log) => log.hookCall.hook_event_name === 'BeforeAgent',
84
+ );
85
+
86
+ expect(beforeAgentLog).toBeDefined();
87
+ expect(beforeAgentLog?.hookCall.stdout).toContain('injected context');
88
+ expect(beforeAgentLog?.hookCall.stdout).toContain('"decision":"allow"');
89
+ expect(beforeAgentLog?.hookCall.stdout).toContain(
90
+ 'SYSTEM INSTRUCTION: This is injected context.',
91
+ );
92
+ });
93
+ });
94
+
95
+ describe('AfterAgent Hooks', () => {
96
+ it('should receive prompt and response in AfterAgent hook', async () => {
97
+ await rig.setup('should receive prompt and response in AfterAgent hook', {
98
+ fakeResponsesPath: join(
99
+ import.meta.dirname,
100
+ 'hooks-agent-flow.responses',
101
+ ),
102
+ });
103
+
104
+ const hookScript = `
105
+ const fs = require('fs');
106
+ try {
107
+ const input = fs.readFileSync(0, 'utf-8');
108
+ console.error('DEBUG: AfterAgent hook input received');
109
+ process.stdout.write("Received Input: " + input);
110
+ } catch (err) {
111
+ console.error('Hook Failed:', err);
112
+ process.exit(1);
113
+ }
114
+ `;
115
+
116
+ const scriptPath = rig.createScript('after_agent_verify.cjs', hookScript);
117
+
118
+ rig.setup('should receive prompt and response in AfterAgent hook', {
119
+ settings: {
120
+ hooksConfig: {
121
+ enabled: true,
122
+ },
123
+ hooks: {
124
+ AfterAgent: [
125
+ {
126
+ hooks: [
127
+ {
128
+ type: 'command',
129
+ command: normalizePath(`node "${scriptPath}"`)!,
130
+ timeout: 5000,
131
+ },
132
+ ],
133
+ },
134
+ ],
135
+ },
136
+ },
137
+ });
138
+
139
+ await rig.run({ args: 'Hello validation' });
140
+
141
+ const hookTelemetryFound = await rig.waitForTelemetryEvent('hook_call');
142
+ expect(hookTelemetryFound).toBeTruthy();
143
+
144
+ const hookLogs = rig.readHookLogs();
145
+ const afterAgentLog = hookLogs.find(
146
+ (log) => log.hookCall.hook_event_name === 'AfterAgent',
147
+ );
148
+
149
+ expect(afterAgentLog).toBeDefined();
150
+ // Verify the hook stdout contains the input we echoed which proves the
151
+ // hook received the prompt and response
152
+ expect(afterAgentLog?.hookCall.stdout).toContain('Received Input');
153
+ expect(afterAgentLog?.hookCall.stdout).toContain('Hello validation');
154
+ // The fake response contains "Hello World"
155
+ expect(afterAgentLog?.hookCall.stdout).toContain('Hello World');
156
+ });
157
+
158
+ it('should process clearContext in AfterAgent hook output', async () => {
159
+ rig.setup('should process clearContext in AfterAgent hook output', {
160
+ fakeResponsesPath: join(
161
+ import.meta.dirname,
162
+ 'hooks-system.after-agent.responses',
163
+ ),
164
+ });
165
+
166
+ // BeforeModel hook to track message counts across LLM calls
167
+ const messageCountFile = join(rig.testDir!, 'message-counts.json');
168
+ const escapedPath = JSON.stringify(messageCountFile);
169
+ const beforeModelScript = `
170
+ const fs = require('fs');
171
+ const input = JSON.parse(fs.readFileSync(0, 'utf-8'));
172
+ const messageCount = input.llm_request?.contents?.length || 0;
173
+ let counts = [];
174
+ try { counts = JSON.parse(fs.readFileSync(${escapedPath}, 'utf-8')); } catch (e) {}
175
+ counts.push(messageCount);
176
+ fs.writeFileSync(${escapedPath}, JSON.stringify(counts));
177
+ console.log(JSON.stringify({ decision: 'allow' }));
178
+ `;
179
+ const beforeModelScriptPath = rig.createScript(
180
+ 'before_model_counter.cjs',
181
+ beforeModelScript,
182
+ );
183
+
184
+ const afterAgentScript = `
185
+ const fs = require('fs');
186
+ const input = JSON.parse(fs.readFileSync(0, 'utf-8'));
187
+ if (input.stop_hook_active) {
188
+ // Retry turn: allow execution to proceed (breaks the loop)
189
+ console.log(JSON.stringify({ decision: 'allow' }));
190
+ } else {
191
+ // First call: block and clear context to trigger the retry
192
+ console.log(JSON.stringify({
193
+ decision: 'block',
194
+ reason: 'Security policy triggered',
195
+ hookSpecificOutput: {
196
+ hookEventName: 'AfterAgent',
197
+ clearContext: true
198
+ }
199
+ }));
200
+ }
201
+ `;
202
+ const afterAgentScriptPath = rig.createScript(
203
+ 'after_agent_clear.cjs',
204
+ afterAgentScript,
205
+ );
206
+
207
+ rig.setup('should process clearContext in AfterAgent hook output', {
208
+ settings: {
209
+ hooksConfig: {
210
+ enabled: true,
211
+ },
212
+ hooks: {
213
+ BeforeModel: [
214
+ {
215
+ hooks: [
216
+ {
217
+ type: 'command',
218
+ command: normalizePath(`node "${beforeModelScriptPath}"`)!,
219
+ timeout: 5000,
220
+ },
221
+ ],
222
+ },
223
+ ],
224
+ AfterAgent: [
225
+ {
226
+ hooks: [
227
+ {
228
+ type: 'command',
229
+ command: normalizePath(`node "${afterAgentScriptPath}"`)!,
230
+ timeout: 5000,
231
+ },
232
+ ],
233
+ },
234
+ ],
235
+ },
236
+ },
237
+ });
238
+
239
+ const result = await rig.run({ args: 'Hello test' });
240
+
241
+ const hookTelemetryFound = await rig.waitForTelemetryEvent('hook_call');
242
+ expect(hookTelemetryFound).toBeTruthy();
243
+
244
+ const hookLogs = rig.readHookLogs();
245
+ const afterAgentLog = hookLogs.find(
246
+ (log) => log.hookCall.hook_event_name === 'AfterAgent',
247
+ );
248
+
249
+ expect(afterAgentLog).toBeDefined();
250
+ expect(afterAgentLog?.hookCall.stdout).toContain('clearContext');
251
+ expect(afterAgentLog?.hookCall.stdout).toContain('true');
252
+ expect(result).toContain('Security policy triggered');
253
+
254
+ // Verify context was cleared: second call should not have more messages than first
255
+ const countsRaw = rig.readFile('message-counts.json');
256
+ const counts = JSON.parse(countsRaw) as number[];
257
+ expect(counts.length).toBeGreaterThanOrEqual(2);
258
+ expect(counts[1]).toBeLessThanOrEqual(counts[0]);
259
+ });
260
+ });
261
+
262
+ describe('Multi-step Loops', () => {
263
+ it('should fire BeforeAgent and AfterAgent exactly once per turn despite tool calls', async () => {
264
+ await rig.setup(
265
+ 'should fire BeforeAgent and AfterAgent exactly once per turn despite tool calls',
266
+ {
267
+ fakeResponsesPath: join(
268
+ import.meta.dirname,
269
+ 'hooks-agent-flow-multistep.responses',
270
+ ),
271
+ },
272
+ );
273
+
274
+ // Create script files for hooks
275
+ const baPath = rig.createScript(
276
+ 'ba_fired.cjs',
277
+ "console.log('BeforeAgent Fired');",
278
+ );
279
+ const aaPath = rig.createScript(
280
+ 'aa_fired.cjs',
281
+ "console.log('AfterAgent Fired');",
282
+ );
283
+
284
+ await rig.setup(
285
+ 'should fire BeforeAgent and AfterAgent exactly once per turn despite tool calls',
286
+ {
287
+ settings: {
288
+ hooksConfig: {
289
+ enabled: true,
290
+ },
291
+ hooks: {
292
+ BeforeAgent: [
293
+ {
294
+ hooks: [
295
+ {
296
+ type: 'command',
297
+ command: normalizePath(`node "${baPath}"`)!,
298
+ timeout: 5000,
299
+ },
300
+ ],
301
+ },
302
+ ],
303
+ AfterAgent: [
304
+ {
305
+ hooks: [
306
+ {
307
+ type: 'command',
308
+ command: normalizePath(`node "${aaPath}"`)!,
309
+ timeout: 5000,
310
+ },
311
+ ],
312
+ },
313
+ ],
314
+ },
315
+ },
316
+ },
317
+ );
318
+
319
+ await rig.run({ args: 'Do a multi-step task' });
320
+
321
+ const hookLogs = rig.readHookLogs();
322
+ const beforeAgentLogs = hookLogs.filter(
323
+ (log) => log.hookCall.hook_event_name === 'BeforeAgent',
324
+ );
325
+ const afterAgentLogs = hookLogs.filter(
326
+ (log) => log.hookCall.hook_event_name === 'AfterAgent',
327
+ );
328
+
329
+ expect(beforeAgentLogs).toHaveLength(1);
330
+
331
+ expect(afterAgentLogs).toHaveLength(1);
332
+
333
+ const afterAgentLog = afterAgentLogs[0];
334
+ expect(afterAgentLog).toBeDefined();
335
+ expect(afterAgentLog?.hookCall.stdout).toContain('AfterAgent Fired');
336
+ });
337
+ });
338
+ });
integration-tests/hooks-system.after-agent.responses ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Hi there!"}],"role":"model"},"finishReason":"STOP","index":0}]}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Clarification: I am a bot."}],"role":"model"},"finishReason":"STOP","index":0}]}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Security policy triggered"}],"role":"model"},"finishReason":"STOP","index":0}]}]}
integration-tests/hooks-system.after-model.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Addressing the Inquiry**\n\nI've grasped the core of the user's question and identified that no tools are needed. My focus is now on crafting a straightforward, direct response that fully addresses their query without any unnecessary complexity. The goal is to provide a clear and concise answer.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12777,"totalTokenCount":12802,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12777}],"thoughtsTokenCount":25}},{"candidates":[{"content":{"parts":[{"text":"4","thoughtSignature":"CiQBcsjafBFqw6veocEvtOGGuQcsyHdcNrXDIn19n9ImwBBwcYQKdgFyyNp8g7o8Ji++OXoqml4gbLPIB2DQbXcaRQfRuYefF8RxMEpzJSITZBlT1VpJQoeYmQcb9c8dg/POmo5d3ZcuLbpVJpbjMIV1SoUI4KEn3zqz7a8BFuyq3zY4VEliRWMZO21JMd8qp59M9m64hX7W1YPyzu8KPwFyyNp8aNCD7P1NJDG3csQkiMW/0jWdPkh+7+XxT7i3ku/lYH4yTEShdicPcmnzoPGhEWTUDr/4Lx+A0DnVGQ=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12777,"totalTokenCount":12802,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12777}],"thoughtsTokenCount":25}}]}
integration-tests/hooks-system.after-tool-context.responses ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Analyzing File Access**\n\nI've realized the `read_file` tool is perfect for accessing the contents of `test-file.txt`. My next step is to call this tool and set the `file_path` parameter to `test-file.txt`.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12785,"totalTokenCount":12841,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12785}],"thoughtsTokenCount":56}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"read_file","args":{"file_path":"test-file.txt"}},"thoughtSignature":"CiQBcsjafE9D7iAF+V3wpXP81/VmxiMeSFA6afML/lAB76U6QFQKXgFyyNp8i/vhxpkTQ5Cq81QTeEJDDMaYihzSTFMqO4Vj0+CLNtoy+SC/LmqA+WaXh4tm6UCNFTzB2fpVW13YOU1oVYhLpVpeck746YExu1MOSTAq7AC9Yz8ZoelXdecKdwFyyNp8q0PejiY9K1osdOJ02tOHAzAb8ZCSFHtHamEPxRB93krGMNvuIYC1jM1JnC/fzpH8gYV+0/xkoPJMHpF/aSzWq4kZ/j5cUhMYaqKJTulY8ZZGfawnXG7z0spmmr06gwfgILa+HK++xQhhTphMQCobX5hyCjUBcsjafHY6eJfVNitYmfruLV1mnoYnNViHuAOOOni9jIz4VMIjLbClKkb2rpVfHIjx+vZSHA=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12785,"candidatesTokenCount":20,"totalTokenCount":12861,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12785}],"thoughtsTokenCount":56}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"This"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12889,"candidatesTokenCount":1,"totalTokenCount":12890,"cachedContentTokenCount":12206,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12889}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12206}]}},{"candidates":[{"content":{"parts":[{"text":" is test content"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12889,"candidatesTokenCount":4,"totalTokenCount":12893,"cachedContentTokenCount":12206,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12889}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12206}]}}]}
integration-tests/hooks-system.allow-tool.responses ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Formulating the Write Operation**\n\nOkay, I'm now clear on the user's intent: they want a file called `approved.txt`, containing the words \"Approved content.\" I've decided to leverage the `write_file` tool. The specific parameter assignments seem straightforward; `file_path` will be \"approved.txt\", and the file's `content` will precisely mirror the desired output string.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12778,"totalTokenCount":12838,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12778}],"thoughtsTokenCount":60}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"content":"Approved content","file_path":"approved.txt"}},"thoughtSignature":"CiQBcsjafF4NswdygCBTU7cA/yXVRcUI3XHwV+E8BDg/hRr1MaoKZQFyyNp8HRY1qEvivtg0LpYPo1022IfTY3QIeigqGvSoRVospxT5MBggc9nRbwH2vrdhZ772IdqOCrpjNHs3wc+h0AF4JzjlBet6+yC2m7TdenVOkzVAtqnNDMQAIS1gDZyKs8w/CngBcsjafOeuyDQtxuK7JCafKjtfvPvoKOkVxzDetQtHesBkPtv1Xng9dkP77jLH44hn9rrg7yA+za6vssiFZUjC/FU25pCWQgIhM+K7nt3wbAgoOZRqra2gRr3od2D3osV/UpYhy8MoloykqrWvHDOzT/0KScpHarwKXQFyyNp8qabyDYlfElywQBjqQT4f6My7+Ln9AbKZQz4NaEe90ESg4jr4jjANxyd/WKzRheaBq7BYxTHQSeShgQbVjk2D0tZO4hAN+CToMtQwJl95Ss4ZEov6gAwMNA=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12778,"candidatesTokenCount":24,"totalTokenCount":12862,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12778}],"thoughtsTokenCount":60}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12939,"totalTokenCount":12939,"cachedContentTokenCount":12203,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12939}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12203}]}}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12939,"totalTokenCount":12939,"cachedContentTokenCount":12203,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12939}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12203}]}}]}
4
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Confirming File Creation**\n\nI've successfully created the file `approved.txt`, and I've verified that it contains the intended content, \"Approved content\". Moving on.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12887,"totalTokenCount":12932,"cachedContentTokenCount":12198,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12887}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12198}],"thoughtsTokenCount":45}},{"candidates":[{"content":{"parts":[{"text":"**Assessing File Contents**\n\nI'm now checking the content of `approved.txt`. I used `cat` to display its contents, and it confirms the initial content of \"Approved content\" is present. My next step will be based on this verification.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12887,"totalTokenCount":12946,"cachedContentTokenCount":12198,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12887}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12198}],"thoughtsTokenCount":59}},{"candidates":[{"content":{"parts":[{"text":"I have created the file. What would you like me to do next?","thoughtSignature":"CiQBcsjafEAq9BWRwBqUousKwXME0A2Wh1tJI5cJC9ROpr9Cix8KagFyyNp81VagWC/YxtY8zCAiThU3BHMVh5wZIsGIWv1NNIXqACLQLoSeLhWEneb6CBkKdbKBugy6g9+jP5phYt+Vz5oYuO1Op2kM1qWjFmEQyr71TUISNtZ9zrOHNQKKW7K9ukUi0paw85YKoAEBcsjafF6QLINjBWwQPZh6EPVNGk4wojTKglNp7xy5vclYBbq58A6A8AtZUHKYA2cV32SLb2TGcPnkE4iKunvPf6sZy9Uc7gKA+x/OgSl7i5m0wSpMOh9fLpGt4CNtieigpxHkNAdxdZ5qzGvCkBFWYhaZAWGbj7+1YibIKJFNjX9yEz1T5dOQmVmceu80dFyz+fwl7RiOXSGR5xK4J7DeClYBcsjafPUccUubdSVLFmRohU4bBtQzLvXxw25mqm5TKANLKINQoloZ+xfXzfe8xw/WZL/mg30AqQErBXPNnLk5vIWLK7suuFAZ7oXdisTCj3MRa1HQmQ=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12887,"candidatesTokenCount":13,"totalTokenCount":12959,"cachedContentTokenCount":12198,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12887}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12198}],"thoughtsTokenCount":59}}]}
integration-tests/hooks-system.before-model.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Initiating String Output**\n\nI'm now fully focused on directly outputting the specified string. The process has been simplified to its core objective, eliminating extraneous steps. All systems are go for immediate execution of the requested string output.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12419,"totalTokenCount":12439,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12419}],"thoughtsTokenCount":20}},{"candidates":[{"content":{"parts":[{"text":"The security hook modified this request successfully.","thoughtSignature":"CiQBcsjafAsmW87n4ndCW3YiNIqK6jp0zaTwTjz12vWiwbCFNAUKdQFyyNp808SX5BqCBNZt+dlgsPf74u9W6ofevKGwkTTHQZWJEQiJR2j4uRfESTazuawuWfzKfNJq5Zml6fokNR9jzmQM+Jf4FHw95Jd4lneap+YGO9x5nZMNDI1cHRx0vs4BYW9GWY7lBIM8xKtaEkPrwqc88goiAXLI2nx5o6VrBpXs6jzf5maZIauSYw42zlnkqdDEMI20rg=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12419,"candidatesTokenCount":7,"totalTokenCount":12446,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12419}],"thoughtsTokenCount":20}}]}
integration-tests/hooks-system.before-tool-stop.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Initializing File Creation**\n\nI'm starting to think about how to make a new file called `test.txt`. My plan is to use a `write_file` tool. I'll need to specify the location and what the file should contain. For now, it will be empty.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":13216,"totalTokenCount":13269,"promptTokensDetails":[{"modality":"TEXT","tokenCount":13216}],"thoughtsTokenCount":53}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"file_path":"test.txt","content":""}},"thoughtSignature":"CiQBcsjafJ20Qbx0YvING6aZ0wYoGWJh3eqornOG4E4AfBLiVsQKXwFyyNp8UlwYs/pv9IRQQGhDlrmlOJF2hfQijryyUYLI+qjDYTpZ6KKIfZF4+vS0soL2BJ3eTXA6gaadFEfNQem3WQVeQoKLFoW4Hv4mbasXqQc0K3p15DuSAtZZENTbCnsBcsjafGK+BJyF/Npnd7gyU0TL5PXePT0nuDFjhJDxlSRUJHDP315TewD3PUYsXd10oWsfhy4B5AngyUiBPUoajdsxg8WxaxnOZYqcp8EIuwtGZrCTev6IihT5nE5jj7u0P9vtnCmkAc6p+4O7Q7Jku1uVGqeJChgzI4YKSAFyyNp8EXSdbttV4xzX+NLKkc276L8Y63tnKU6/Y7fc9/58tU29DSdrgwfe9qmvwtTsO0piFXSLazqHJt8h2bgR7A7GnKDiIA=="}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":13216,"candidatesTokenCount":21,"totalTokenCount":13290,"promptTokensDetails":[{"modality":"TEXT","tokenCount":13216}],"thoughtsTokenCount":53}},{"candidates":[{"content":{"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":13216,"candidatesTokenCount":21,"totalTokenCount":13290,"promptTokensDetails":[{"modality":"TEXT","tokenCount":13216}],"thoughtsTokenCount":53}}]}
integration-tests/hooks-system.compress-auto.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Devising a Greeting Phrase**\n\nI've been occupied by the constraint of constructing a five-word salutation. My goal is to make it natural and concise. I'm exploring various combinations to meet the specified word count precisely.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12587,"totalTokenCount":12612,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12587}],"thoughtsTokenCount":25}},{"candidates":[{"content":{"parts":[{"text":"Hello! How can I help you?","thoughtSignature":"CiQBcsjafHso9FUsdYOCTv1xOLlW4MnjbeYnUUBocz0KNgHSzOcKZAFyyNp8XuI6j2afRczgPL8v1dxfVwAJ+5XDKhWKIYf1/8TKGVHh7xXnPfdYBdQ07Ohe7OZXr92xL/IC7B1U2SHDuAOozC0CCW7aiDysu6Hbo6jzYfW5epKht4QjdxYgcKHySrkKMQFyyNp8jXWlHmox53O/CJPXXz2FAmw+ubHKBpYgRezBpA+byyEY2RbVYlZlEMSNkhs="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12587,"candidatesTokenCount":7,"totalTokenCount":12619,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12587}],"thoughtsTokenCount":25}}]}
integration-tests/hooks-system.input-modification.responses ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"content":"original content","file_path":"original.txt"}}}],"role":"model"},"finishReason":"STOP","index":0}]}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"I have created the file."}],"role":"model"},"finishReason":"STOP","index":0}]}]}
integration-tests/hooks-system.input-validation.responses ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Defining File Creation**\n\nI'm thinking about the user's intent to generate a file named \"input-test.txt\" with the content \" test\". I've determined that the `write_file` tool is suitable. I've parsed `file_path` as \"input-test.txt\" and `content` as \" test\". This should accomplish the user's need.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12778,"totalTokenCount":12840,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12778}],"thoughtsTokenCount":62}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"file_path":"input-test.txt","content":"test"}},"thoughtSignature":"CiQBcsjafCO/Ifs3Lj/Gtzy2ylSYoGB3GXjJby4F3R8FxWp+hP0KZwFyyNp8oD7KvcSYXDimGOiqAdxtdOJpc2tFJbHm2Jw7ahiuKLtoKZWE+1bBZEWVKxC0dCQIeIcxZ0SaLn7tDbfc2qPzhyUA46d/T1+e314SFLWW1asIOBkQ4T0sFDAFPZ4m9bFm3UkKbAFyyNp8EAnclI0wYCGwpg0AOOV52F5J9Hc2EeaXkGsc6hCnba7aNhPucWYIn2Da8FK2IJAWUWaNvGNGoNUZETaG+iL9+6KRJgN3Ql/wQzQ2pHUvTGHC3RkfMGTQ+YCQKvlOReilps5lDmMnhQpTAXLI2nzcl9Aqd0Nb/w934w+tqz1Jth7GlQVMYktHOl7Hgkoykfh3NzM67SEAilxjowfBL6MY7UBUP3YGwi1CXVVa4d0wHnMD9BJYp2w8ztZch8I="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12778,"candidatesTokenCount":25,"totalTokenCount":12865,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12778}],"thoughtsTokenCount":62}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**File Creation Achieved**\n\nI've successfully created the file as requested. Now, I'm ready to move on to the next instruction whenever it arrives. I am now awaiting the next task.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12940,"totalTokenCount":12965,"cachedContentTokenCount":12203,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12940}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12203}],"thoughtsTokenCount":25}},{"candidates":[{"content":{"parts":[{"text":"Done.","thoughtSignature":"CiQBcsjafEwfH5zTnAjEjloMcDDflS/MmoH03HXVl8HoQ04vmVIKcQFyyNp8/6HrBz8vokXB1Ms1zW51p32T3Ni3HEbgSFPHMGZt9LHFtLkLzuFrxym66z1Tcb5tqj+7jAdpM/dIUb6ecrKj9FWqMB+QR4BSxdAiJSiL8Rp+Pc5ckCtT1nrv4C5w3/fhCNE4WvZzeyGPt+PACjsBcsjafNWzUJcHxgKp6MYWQ8RW0QrGerM51nkgXHBafxY5KwTznX4B/ETccGnXX3zSciaJiZR1FfudVw=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12940,"candidatesTokenCount":1,"totalTokenCount":12966,"cachedContentTokenCount":12203,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12940}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12203}],"thoughtsTokenCount":25}}]}
integration-tests/hooks-system.multiple-events.responses ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Formulating a Plan**\n\nOkay, I've outlined the initial steps: I'll use the `write_file` tool to make a file named `multi-event-test.txt` containing the text \"testing multiple events\". After that, I'll need to remember to reply with the phrase as requested. It seems straightforward so far.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12622,"totalTokenCount":12692,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12622}],"thoughtsTokenCount":70}},{"candidates":[{"content":{"parts":[{"text":"**Confirming the Procedure**\n\nI've solidified the steps. First, I'll create `multi-event-test.txt` using the `write_file` tool with the required content. Following that, my response will be \"BeforeAgent: User request processed.\" This ensures I fulfill both parts of the request.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12622,"totalTokenCount":12713,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12622}],"thoughtsTokenCount":91}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"file_path":"multi-event-test.txt","content":"testing multiple events"}},"thoughtSignature":"CiQBcsjafIqcYtNLIeBwJi3k5k8jho3QiWM+51Kw5vTQ7/V4qVQKZgFyyNp8mIIB0+Mvwhvo2fACDpTWpRYeOFPGrjZrc+N05S0WGEHzE4Dv9peHKdvZkjGNW+HyYHXoRpd5c/ScdhPxQoVZmZ9K7sRjVxv/nWVDoKnHlSsn94nJ8acjLnj1oqt9cHni0ApyAXLI2nwj5WuLHr+UFIxnqRKCUJboLo6bQMkqR1TsqXbjsgHp3zNQYT+xzbse4PKPLJV48FN6cL9MrrZ81E7k7AVo1cKyrC7ky7tdRH6gYHewIqgQWBIUgMKhLkePH/fYZ6fS7SMrf4Q6DFGHh6pIAAdRCooBAXLI2nxpudEZr+5jZAaAcCMIdij5oZq3s0xsQv/7iWVh8IossRuR0J4eMMSN8fV6+fjbSQ6YtJQfrxsm3a6gVIkJNno2b2PRZestS/0Z7DvPDGE6r1sGchvbcz8EW7Z/pvJvPBRFWlMTJ1eqY9vuyuNYMKeWlyt+5V9y2GUbcLWvcNDZSC43vQEKCo0BAXLI2nxP4INgBaSHInyFrG1/SEP0SUimKvP69FkcIBxx60x3iKqdtb2flLIhoOr/QuesASlflRfzNo3J5LOudrjZzNlRfVRqOZIyOVxZlviXtO7+w/oPCV61Sby6xPTGtFsWlt6GxEGF7iYLfvi4KWN9q/W9tlqEqUrpl/WMwS/4pYBi1xPcvXZNlJ6g"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12622,"candidatesTokenCount":28,"totalTokenCount":12741,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12622}],"thoughtsTokenCount":91}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12836,"totalTokenCount":12836,"cachedContentTokenCount":12204,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12836}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12204}]}},{"candidates":[{"content":{"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12836,"totalTokenCount":12836,"cachedContentTokenCount":12204,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12836}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12204}]}}]}
3
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"\n"}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12836,"totalTokenCount":12836,"cachedContentTokenCount":12204,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12836}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12204}]}},{"candidates":[{"content":{"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12836,"totalTokenCount":12836,"cachedContentTokenCount":12204,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12836}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12204}]}}]}
4
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Echoing User Commands**\n\nI'm now tasked with echoing a specific phrase following a particular signal, but it's becoming complex. The user wants me to repeat \"BeforeAgent: User request processed\" when prompted. It appears I need to retain context from the previous turn, the user's initial request to create a file, to correctly respond now.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12759,"totalTokenCount":12827,"cachedContentTokenCount":12199,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12759}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12199}],"thoughtsTokenCount":68}},{"candidates":[{"content":{"parts":[{"text":"**Responding Precisely to Prompt**\n\nI've determined I need to repeat the phrase \"BeforeAgent: User request processed,\" even though the overall context and turn history are complex. The user has given several prompts, but has now provided a more direct command, which I believe is to follow up on the previous request. I am taking care to match the specific instructions the user provided.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12759,"totalTokenCount":12982,"cachedContentTokenCount":12199,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12759}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12199}],"thoughtsTokenCount":223}},{"candidates":[{"content":{"parts":[{"text":"BeforeAgent: User request processed","thoughtSignature":"CiQBcsjafAntJrb1JBgpnZaCNeYhOJXtbH6dKTeM1llglCdoOvUKYwFyyNp8PUj5sihYyITQJhdz4MqEeftyuUc4G+iTprve11gPN04eK9Y1Wi/wyln4RjRgroIrV5kByKzdGhECoyCeInpiILGhY0peIM7dZOKFdIOL7xAR9pmn4wMreqyH7l5WSAqJAQFyyNp8Cugemkt4YZWkIwEJYmUukLFx4d5EwP/9k/e4OH/svpM+uyuN3n1KVN3bFgRV5yuF0HnDLl+P7WVSSxMmWvXO2f7A1HALg+gCvZw9IV7Btgg1qp81dDoNcVkzSbTBtT4UrlJ5R6sclvHZOLUtKGwBEQ6zRonBugAgj9RV4BT1AJNOgdSsCokBAXLI2nyDGU1Iq30QVbqhgEwFa5sB6uPC+35BV8ZKGwK+YglO9rqXMrkXM+GcQi2hVIsOFXBYGTS6E2/mQfFbIKDytrb1JgP3q5xVd/bE23M2Nnf+q5TLbRpLAPmyfg0AGwhN0L7d5W6b/3ydqEPeA1/Vw/cnBzz5ND1LOTOX6BFqEs33/WHj7HIKpAEBcsjafEsn8//cZMWUQcSAucBQauojv/f7h11nbeMrZK84nEotR30BgMIWYiiWM6sGDy/4MzHwr+z2YdAz4PSgRvEf7DPxHps2nvZfAdtskgtdPl2JD81WpokSnJvCqU+cOuz+Nh3+fIiZ6vEsVpi/5cwEiGT0g3Z3I2ubyzv58oH8YnVQlKT3MsKRGb5//aXZJY57jNrexgDPzYAQsBgSuGBmqwqaAQFyyNp8sSIYw3It6GpZqC+oxJCC26pt4RxhG8rDZ3zuoADYlOpoUdSzbNuDB+iVHeen5OoCEAaH0GrFV4iZxgu40wu4ZD/VMfHi/Vm7vku23EUV/94U8mT+VEwPfd2gqv+3xPZ9MEHjOOox1Xq1984w2cA6u0Qn7wWHXeOGFVGSOHtdJtQ7ToNT8VEecblAVq8lm42sSccXQEEKmAEBcsjafONCvBhW2s8Bset20YFdbeSHelnILFDxXlCoYla5nP5UjGk4vpXu2+7RCFtKXfoyYEVEkmiGBRsmwJ82Q1nMkGkXMhuTdNhu4aCwI5m+STGxx26vkp9bcqGwMDHBotZL63PSrJacRoW8zfpDXD1PABLeTIfh5jgipQdgltyjlbc+3qfIfjBYNRSkE8ByErSz5rT7SwqSAQFyyNp8W2kut1PSJISxM7YJtbRdFqPBTikGDM6F/3l6ba6LpeRBfHdtueLChqFpwLH41VdIPQ7lRZflOq3KaZz+TQ11eDnYQbiaIdGOPgHJ/HH/0iQv2hnoOY5vg3gubFWFuZh9Bfun2VCYUI39tIxGC46TZWfgCdiP/O9CFOlpDfidPiz5ZS/4LhG9FA4Q85OuCpEBAXLI2nzpoEUA6jCZopeNTRA2uZ1r0DMm5cWVVXtFO4CoRS+19BbADNBRyNrR5qcf7bUflJBvMRVxx3mtmgK9aE5VmKYxK2Dqg15l9RUxjtqspC3VVmszVd6lOkf1BBQ/VtWDulqRetKE2u62Is9NNGuK9HsLzIBLRRc8QoML41WffuXQ+uxwyXpjx2USC44MGAqIAQFyyNp8gN3lOyHyk674W3Pyv+Egw1ZDUQK4xpvAfgnK+y53gclMGJ2IjOSvg4j0f1WO1OGqY2TBUFS7w21PXasvCkfxpqeStEb+U7Vm0r63LzXdGdug5/b1Ap6Phn4/vAYmfaKISKG4+QpjI+ehgEJzsIee2rgqOaePTP18fq8T7EDbF/B/iscKNQFyyNp8DWt2a8OetaCc5E/KsntbbOcNc7yikPZBdUezphrqIH4ztpicsHvEicYF002qWHoY"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12759,"candidatesTokenCount":4,"totalTokenCount":12986,"cachedContentTokenCount":12199,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12759}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12199}],"thoughtsTokenCount":223}}]}
integration-tests/hooks-system.notification.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Executing the Command**\n\nI've got the command, \"echo test,\" ready to go. My focus is entirely on calling the `run_shell_command` tool now. The user's input is processed, and the next step is straightforward: using the tool to execute the supplied command.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12751,"totalTokenCount":12801,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12751}],"thoughtsTokenCount":50}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"run_shell_command","args":{"command":"echo test","description":"Running the command 'echo test'"}},"thoughtSignature":"CiQBcsjafL1lDlnUGmt38n1/gjwecXzy9S3qEW5sYMEno5Mr7LEKZgFyyNp8jMABmMAatt49FTdh7UiM62SI1GnjcyG+kV7xzcD73uMKHST/0D0vKP7x1equv5d6YiXnOslhVnnHotYPtVl0/kI/0unBZRdMzkBNrJXKUoSWXJXxNpV6JhJav3Uh9h1sPQqOAQFyyNp8PFeESLk0J5cPFP0EA7a13iA/rXTiKoHnjSCzDV9ALcXM78xv10/V028ZtDeQslYfT82q4++W8AlJwTQRTIrdscu2y+nCS8jnQizYN1V1yR42eMzuBU3txXcqEV8bmP6GGOe58vrqyS2zdnJKCgMntMB/niwlJlr5frhDestSOJk62tVDWKFzOiAKOAFyyNp81FtGXQTX+OSio/2PbzpCCuaQFqpEgCZpkaXXyvmXYDAI1qCq1tA+m/e5ozWdm8zTGuyb"}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12751,"candidatesTokenCount":28,"totalTokenCount":12829,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12751}],"thoughtsTokenCount":50}}]}
integration-tests/hooks-system.sequential-execution.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Seeking Task Clarity**\n\nI'm currently focused on identifying the precise task. My initial assessment indicates the user is seeking assistance, but the specific requirements remain undefined. I will directly solicit a detailed task description from the user to clarify this.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12604,"totalTokenCount":12633,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12604}],"thoughtsTokenCount":29}},{"candidates":[{"content":{"parts":[{"text":"Hello! I'm ready to help. Please describe the task you'd like me to assist you with.","thoughtSignature":"CiQBcsjafM2CL00L595T19DK8M8zP5p9/tbFPPwdM2S6669z2FgKYQFyyNp8Ya0YVCtft9Asr/45XOCfNdPWbwZt8SvIeX3IxYzOFcOK14+DnoDIuTIrmRQBeUvdxD59QmEWx+/OaSxj9564L0IU703C1JX20buEtYhkRM4LhK0G4LG/z6IJauEKSQFyyNp8n784BnEcDTQGfZ8/s3pl/TNaNzjQx0o8wYCYZH1qsRbVa3YJAvRGrVXL6y9ka10w0lhEsrQ8vOiw6ilZKirA5DjLz4U="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12604,"candidatesTokenCount":22,"totalTokenCount":12655,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12604}],"thoughtsTokenCount":29}}]}
integration-tests/hooks-system.session-clear.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Greeting the User**\n\nI've registered the user's greeting. I'm primed to respond with a friendly welcome and signal my availability to assist. My focus now is drafting a suitable response.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12761,"totalTokenCount":12787,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12761}],"thoughtsTokenCount":26}},{"candidates":[{"content":{"parts":[{"text":"Hello! I'm ready to help. What can I do for you?","thoughtSignature":"CikBcsjafBz/0rqJuIv9woxRvivjZyAqBjpoJhOTSPfcbMWCawTfcyKImQpxAXLI2nxyuBo6dqZmTxkH7XxPxjq7mNoacRa48wc/eT5caK/4tu0Y9fJ1ScpJZb+tCNzrqTNwVXa98ppjB2O/X4eejJN+hUr3LCalDFRdRLO17PFUI5qgYSbSgIGzhbnQASgzOArvvqzDPPgqXWVIDj8KMQFyyNp8ayfqBNRkBykRSTDtzOKVGkjLW1dXWamLB4ojeEVHSOgne4vlYaKs44pitsg="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12761,"candidatesTokenCount":15,"totalTokenCount":12802,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12761}],"thoughtsTokenCount":26}}]}
integration-tests/hooks-system.session-startup.responses ADDED
@@ -0,0 +1 @@
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Initiating a Dialogue**\n\nI've successfully received and understood the user's initial request. My next move will be to output a simple \"Hello\" as a greeting, fulfilling the basic instruction I was given. This constitutes the first step in the interaction, and I'm ready to move forward based on the user's subsequent input.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12588,"totalTokenCount":12607,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12588}],"thoughtsTokenCount":19}},{"candidates":[{"content":{"parts":[{"text":"Hello","thoughtSignature":"CikBcsjafB9jXawgyqQ5mpEJ4ihpLD/B2i8GR75sod00ZF3TCbrLHS9YjgpeAXLI2nx1fmJO2VIiwBpF+vLBPhYE/B2992PVW6XM20cEYx4g0leDNs6BIhzEipm6RYOxzgz8KxH9+ZkCnd8bVZr59lbDCgqSCSB6IKA+csXHKsF9g3UMRAtoSBwiBw=="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12588,"totalTokenCount":12607,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12588}],"thoughtsTokenCount":19}}]}
integration-tests/hooks-system.tail-tool-call.responses ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"functionCall":{"name":"read_file","args":{"file_path":"original.txt"}}}],"role":"model"},"finishReason":"STOP","index":0}]}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"Tail call completed successfully."}],"role":"model"},"finishReason":"STOP","index":0}]}]}
integration-tests/hooks-system.telemetry.responses ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"**Initializing File Creation**\n\nI've decided on the `write_file` tool to create the telemetry file. I'll pass \"telemetry-test.txt\" as the file path, and an empty string for the content, as the user didn't specify anything to include. This is the initial setup; the file should now exist.\n\n\n","thought":true}],"role":"model"},"index":0}],"usageMetadata":{"promptTokenCount":12779,"totalTokenCount":12850,"cachedContentTokenCount":12204,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12779}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12204}],"thoughtsTokenCount":71}},{"candidates":[{"content":{"parts":[{"functionCall":{"name":"write_file","args":{"content":"","file_path":"telemetry-test.txt"}},"thoughtSignature":"CiQBcsjafG+JDSqtKOK+ZvSjZQZmS91c1Gz0YyTiirI2u5+rhEIKZAFyyNp8DNXb+xHILTC+FVlEifqEHdrmfFNBLKojci1UIBhcZpQ4UXCMkxUXYKO34IjTlyLgSsjVbbXWEFXatb/z/RtTDcf51uc3YOEwlDScGempkJxfFgcPfIiD7bhuHBqdQfUKfAFyyNp8wZ71h+QjdfVw12PwDXWgGZ0Xed1GuyJXuqAwpWnwxDIvsDaPwDFYyLR1XDiIZZk4AvFCGt6HGMSLRuPh4K3i9CVnDc5hcjyvMIde0idAFMrgs2Mq5SARfCPrWkqyq2f0Q0WonUl2n7yr/sDQ78rx2E6qXyUJ8XMKfAFyyNp8DdTYLttyI0jknqAeZDxdFmHtpJUI8UKP5YHzpQc8Qn80OJcwhZSRH4HRKCqoC7Sukq/A5vJ5T468WqgjOoLlPLq02bYRTf/q6LC1ogEhdLHrcFv2jDeCdXJJ8NHv3O4DZAUAk1W5Gd0428zMFOxH3AgkWwEGuow="}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12779,"candidatesTokenCount":24,"totalTokenCount":12874,"cachedContentTokenCount":12204,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12779}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12204}],"thoughtsTokenCount":71}}]}
2
+ {"method":"generateContentStream","response":[{"candidates":[{"content":{"parts":[{"text":"OK."}],"role":"model"},"finishReason":"STOP","index":0}],"usageMetadata":{"promptTokenCount":12951,"candidatesTokenCount":2,"totalTokenCount":12953,"cachedContentTokenCount":12202,"promptTokensDetails":[{"modality":"TEXT","tokenCount":12951}],"cacheTokensDetails":[{"modality":"TEXT","tokenCount":12202}]}}]}