| Age | Commit message (Collapse) | Author |
|
Given the approximate nature of the evaluations and truncated tree
searches, the result of negamax will always be sensitive to the
particulars of the depth bounding and the conditions for referring to
precomputed values.
The tradeoff here is a slight performance penalty (~200ms at depth 6
on my old Intel(R) Core(TM) i7-3520M CPU @ 2.90GHz), but a greatly
increased opponent strength as compared to `>=`, and the old
convolutional network (flawed as the implementation was).
|
|
Previously the code base assumed that there was a single, global game
state which was the implicit target of all actions taken. Looking
ahead at architectural improvements, this has now been (almost
entirely) made explicit and functions take tak_state_p where
necessary (and also where unnecessary).
Two important fixes to actions.c were made:
- Previously when generating the possible stack moves, stack height
overflows (> 15) were not taken into account and this resulted in the
tree search corrupting the board state. Now action search does not
list all legal actions, rather the subset of these encodeable by the
implementation.
- The check for crushing on a stack move was incorrect (too strict),
and this resulted in many legitimate moves being igonored.
Finally, in other changes, weights have also been improved by training
all games instead of some subset for chosen players, and clang-format
was run on the codebase.
|
|
Gone is the convolutional neural network, for it turns out not only is
it more difficult to train, but all of the extra information about
board layers didn't make much of a difference at this size.
So cnn1986 has been replaced by nn1986, a standard, two-layer, dense
nn configured as a binary classifier and (mis)used in that capacity.
Note: total number of parameters is unchanged.
HARK: this new nn exposes a bug somewhere in ctak. Run ctlm with
self-play to see the completely borked board state at the end.
|
|
I tried the following, but they all made things worse:
- moving away from the singly-linked (tail tracking) list for actions
by:
+ using an array zipper for a deque
+ using an array to poorly hold a floating deque
- caching the results of generating move lists in the transposition
table and then
+ copying the resulting list/zip/deque instead of generating it
+ applying the move-to-front without copying, but this made the
search order worse. Presumably in this case shallower nodes were
messing up the search tree with garbage moves?
I think some of this is not supposed to happen, but i have just the
right combination of poor evaluation function and naively ordered and
cheap move generation that i'm in a local minimum here.
|
|
Eventually there'll be a more complicated data generation step than
the one we're presently using, so having it in-lined in the loop is
wasteful. Ideally also this would be update per ply and we could avoid
recalculating it entirely for every query -- though it's probably
``fast enough'' for now. Also, caching is WIP.
|
|
|
|
If TEI is implemented, then i could make use of Morten's
racetrack (https://github.com/MortenLohne/racetrack) and develop a
quantitative measure of the bot's performance. This is the current
priority.
|
|
``Common wisdom'' dictates that placements are often better than stack
moves, so we bias the generated move list in this fashion. Seems to be
a little faster.
|
|
|
|
|
|
For now we'll stay with directly recomputing it at each non-terminal
node
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|