Skip to main content
Agent Action is Komos’s most powerful node type. It’s a prompt-driven AI agent that can perform complex, multi-step operations in a single node.

Capabilities

Agent Action can:

When to use Agent Action

Use Agent Action when:
  • The task requires judgment or adaptation
  • Multiple steps need to happen based on what’s found
  • You need to combine browser automation with data processing
  • The exact steps depend on page content
Use specialized nodes when:
  • You need pure data transformation without browser context (Process Data)
  • You want to make a direct HTTP call (API Request)
  • You need to branch or loop (If/Else, Loop)

Creating an Agent Action node

  1. In the builder, click + Add Node and select Agent Action
  2. Write a prompt describing what the agent should do
  3. Define outputs the agent should produce
  4. Optionally attach skills and integrations

Writing effective prompts

Be specific about the goal and constraints:
Tips:
  • Break complex tasks into numbered steps
  • Specify what success looks like
  • Mention edge cases or constraints
  • Don’t over-specify - let the agent adapt

Defining outputs

Declare what variables the agent should produce: Outputs are available as ${node_id.output_name} in downstream nodes. For object and object_list types, you can add an optional schema field listing the expected fields (each with name and type). This helps downstream nodes and API callers understand the shape of the data.

Advanced features

Network request inspection

When data isn’t visible in the DOM (hidden IDs, URLs constructed by JavaScript), use network inspection:
The agent has access to:
  • get_network_requests(url_pattern, method) - Find captured requests
  • get_network_response_body(request_id) - Get response content

Attaching skills

Skills provide reusable guidance. Attach them in the node editor:
  1. Open the Agent Action node
  2. Find the Skills section
  3. Select skills that apply to this task
The agent reads skill instructions alongside your prompt.

Toolsets

The toolsets parameter controls which categories of tools the agent can use during execution. When omitted, the agent has access to all available tool categories. Specify a list of toolset names to restrict the agent to only those capabilities.

Using integrations

Enable connected integrations for API work:
  1. In the node editor, find Integrations
  2. Select which connected accounts to allow
  3. Describe the integration work in your prompt

Visual references

Add screenshots to show the agent what to expect:
  1. Expand Advanced > Visual References
  2. Upload or paste screenshots
  3. Label them clearly (e.g., “Login page”, “Success state”)
Visual references improve reliability for complex UIs.

Combining with other nodes

Agent Action works best as part of a larger flow:
  • Use Login to authenticate with stored credentials
  • Use Agent Action for browser automation, data extraction, and adaptive steps
  • Use Process Data to transform or reformat captured variables
  • Use File Output or Email Send to deliver results

Troubleshooting

Agent doesn’t find elements

  • Add visual references showing the expected UI
  • Be more specific about element descriptions
  • Mention in the prompt that the agent should wait for dynamic content to load

Outputs are empty

  • Verify output names match what’s declared
  • Check the run logs for what the agent captured
  • Ensure the agent prompt asks for the data explicitly

Integration calls fail

  • Verify the integration is connected in Settings > Integrations
  • Check that the task has the integration enabled
  • Review the integration’s required scopes/permissions

Best practices

  1. Start simple: Write a basic prompt, run it, then refine
  2. Use skills for patterns: Extract repeated guidance into skills
  3. Declare all outputs: Don’t rely on implicit variables
  4. Add verification: Include verification steps in the prompt or use Process Data to validate outputs
  5. Check run logs: The execution view shows exactly what the agent did