5 repos
Vision language models and autonomous agents that interact with graphical user interfaces to automate browser and desktop tasks. This cluster covers frameworks, implementations, and models for building computer-use agents—systems that can perceive and act within visual environments like web browsers and operating systems. The central repos (CogAgent, Fara variants, UI-TARS-desktop) represent different approaches to VLM-based agents and their deployment, while the broader collection includes supporting tools, datasets, and agent infrastructure across Python, TypeScript, and Rust.