Did you watch Jev play pokemon? It showed interesting patterns with that game too. I'm going the llm route and having it play my first RPG. Turn based so latency is acceptable, prefer the smarts for this one.
I've been pondering if I even bother with map views, or just provide navigable locations as aist like Jev-Pomemon. It'll end up having both options I'm sure, since I'm using the project to explore harness evals.
since mine is more of a harness/eval framework, I'm going to experiment with different toolsets and context constructions
It's definitely easier to skip that part and just use A* for navigation, but I also want to see it explore and learn the map, through more basic moves. Right now I'm not exposing direct button presses to the agent. Similar for menus, give a list of actions (actor->fight->target) and an algo converts that to button presses, so each turn in a battle, it produces an action for each party member.
The danger is, that you end up passing the decision model something where the choice is trivial. There's so much in the harness that you don't need any intelligence.
I have yet to see and credible accounts of these "system one" models being better than alternatives in the range from [task specific classifier ... llm]
I've seen way more about how brittle and inconsistent Jev is with decisions that should be straight forward, like from pokemon...
- I'd like to buy a potion!
- Are you sure?
- No...
- I'd like to buy a potion!
- Are you sure?
- No...
- <repeat for 5 minutes>
it took it waaaay too many tries to buy item
Then there are the examples like coin flipping or dice rolling, Jev doesn't understand basic probability, picked heads like 84% of the time or something...
Did you watch Jev play pokemon? It showed interesting patterns with that game too. I'm going the llm route and having it play my first RPG. Turn based so latency is acceptable, prefer the smarts for this one.
I've been pondering if I even bother with map views, or just provide navigable locations as aist like Jev-Pomemon. It'll end up having both options I'm sure, since I'm using the project to explore harness evals.
Interesting idea - so instead of "which move do you want to make?" it's "which location do you want to go to?".
yeah, that was how jev-pokemon worked
since mine is more of a harness/eval framework, I'm going to experiment with different toolsets and context constructions
It's definitely easier to skip that part and just use A* for navigation, but I also want to see it explore and learn the map, through more basic moves. Right now I'm not exposing direct button presses to the agent. Similar for menus, give a list of actions (actor->fight->target) and an algo converts that to button presses, so each turn in a battle, it produces an action for each party member.
The danger is, that you end up passing the decision model something where the choice is trivial. There's so much in the harness that you don't need any intelligence.
I have yet to see and credible accounts of these "system one" models being better than alternatives in the range from [task specific classifier ... llm]
I've seen way more about how brittle and inconsistent Jev is with decisions that should be straight forward, like from pokemon...
- I'd like to buy a potion!
- Are you sure?
- No...
- I'd like to buy a potion!
- Are you sure?
- No...
- <repeat for 5 minutes>
it took it waaaay too many tries to buy item
Then there are the examples like coin flipping or dice rolling, Jev doesn't understand basic probability, picked heads like 84% of the time or something...