Was curious how LLMs would play chicken without a clear reward matrix, and then took it a step further by allowing communication and realtime decision making.
Each LLM is in an independent loop, where each turn they are given the current speed and time to impact, as well as their historical latency. They can pre-plan moves to account for latency.
Each LLM is in an independent loop, where each turn they are given the current speed and time to impact, as well as their historical latency. They can pre-plan moves to account for latency.