aboutsummaryrefslogtreecommitdiff
path: root/include/negamax.c
AgeCommit message (Collapse)Author
41 hoursfix longstanding negamax bugHEADmaintslil
we needed to be saving the alpha before the search loop in which we modified it. At this point, though i can't remember, i expect that this is the cause of whatever "instability" the comment was talking about that would have prevented us from using >= in the TT lookup.
41 hoursswitch to returning an array of actions instead of a linked listtslil clingman
- attempts to keep the same move ordering as the list method - saves ~170msec on a depth 7 search for a given board configuration - there is room to improve the pre-allocation size estimates, these bounds are not obviously tight and it may or may not be faster to have tighter bounds or even some form of estimation
41 hoursmight as well enable size 6tslil clingman
Same network architecture, same training principle. Predictably this is too slow. Also statically allocate state in driver programmes.
41 hoursChange transposition table lookup policytslil clingman
Given the approximate nature of the evaluations and truncated tree searches, the result of negamax will always be sensitive to the particulars of the depth bounding and the conditions for referring to precomputed values. The tradeoff here is a slight performance penalty (~200ms at depth 6 on my old Intel(R) Core(TM) i7-3520M CPU @ 2.90GHz), but a greatly increased opponent strength as compared to `>=`, and the old convolutional network (flawed as the implementation was).
41 hoursswitch to explicit game state & important bug fix & clang formattslil clingman
Previously the code base assumed that there was a single, global game state which was the implicit target of all actions taken. Looking ahead at architectural improvements, this has now been (almost entirely) made explicit and functions take tak_state_p where necessary (and also where unnecessary). Two important fixes to actions.c were made: - Previously when generating the possible stack moves, stack height overflows (> 15) were not taken into account and this resulted in the tree search corrupting the board state. Now action search does not list all legal actions, rather the subset of these encodeable by the implementation. - The check for crushing on a stack move was incorrect (too strict), and this resulted in many legitimate moves being igonored. Finally, in other changes, weights have also been improved by training all games instead of some subset for chosen players, and clang-format was run on the codebase.
41 hoursnew neural network arch (faster + better) & minor changes + fixestslil clingman
Gone is the convolutional neural network, for it turns out not only is it more difficult to train, but all of the extra information about board layers didn't make much of a difference at this size. So cnn1986 has been replaced by nn1986, a standard, two-layer, dense nn configured as a binary classifier and (mis)used in that capacity. Note: total number of parameters is unchanged. HARK: this new nn exposes a bug somewhere in ctak. Run ctlm with self-play to see the completely borked board state at the end.
41 hoursTrying to make things fastertslil
I tried the following, but they all made things worse: - moving away from the singly-linked (tail tracking) list for actions by: + using an array zipper for a deque + using an array to poorly hold a floating deque - caching the results of generating move lists in the transposition table and then + copying the resulting list/zip/deque instead of generating it + applying the move-to-front without copying, but this made the search order worse. Presumably in this case shallower nodes were messing up the search tree with garbage moves? I think some of this is not supposed to happen, but i have just the right combination of poor evaluation function and naively ordered and cheap move generation that i'm in a local minimum here.
41 hoursFix copyright notice in files, and small preemptive optimisationtslil clingman
Eventually there'll be a more complicated data generation step than the one we're presently using, so having it in-lined in the loop is wasteful. Ideally also this would be update per ply and we could avoid recalculating it entirely for every query -- though it's probably ``fast enough'' for now. Also, caching is WIP.
41 hoursRename ct_k -> ct, IANAL but ...tslil
41 hoursTried some naive iterative deepening. Work on TEI interface nexttslil clingman
If TEI is implemented, then i could make use of Morten's racetrack (https://github.com/MortenLohne/racetrack) and develop a quantitative measure of the bot's performance. This is the current priority.
41 hoursTry to squeeze out a little more performancetslil clingman
``Common wisdom'' dictates that placements are often better than stack moves, so we bias the generated move list in this fashion. Seems to be a little faster.
41 hoursAdded license information!tslil clingman
41 hoursPrint progress before recursion rather than aftertslil clingman
41 hoursIt would appear that any function call whatsoever is slower :/tslil clingman
For now we'll stay with directly recomputing it at each non-terminal node
41 hoursI don't have the presence of mind to debug this right nowtslil clingman
41 hoursSmall changes to build ct1986tslil clingman
41 hoursStoring best moves!tslil clingman
41 hoursThis matches alpha-beta!tslil clingman
41 hoursStable negamax-alpha-beta fail-softtslil clingman
41 hoursOnce again, adding TT changes the outcometslil clingman
41 hoursActually it seems before i was mis-countingtslil clingman
41 hoursThere is still a bug, it doesn't appear to be checking enoughtslil clingman
41 hoursStripping debug stufftslil clingman
41 hoursStill bugs...tslil clingman
41 hoursStill some bugs, standing stone becomes flat at depth4 self-play??tslil clingman
41 hoursThere are still some bugs in the undo almost surely...tslil clingman
41 hoursLots of bugfixes, mostly uint vs int. Still weirdness in gametslil clingman
41 hoursThis is the basic idea, there's ≥ 1 bug (generates illegals...)tslil clingman
41 hoursWorking on action lists to refactortslil clingman
41 hoursStill tryingtslil clingman
|
41 hoursHash collisionstslil clingman
41 hoursFixed some important bugstslil clingman
41 hoursAttempting Zobrist hashingtslil clingman