aboutsummaryrefslogtreecommitdiff
path: root/include
AgeCommit message (Collapse)Author
42 hoursswitch to explicit game state & important bug fix & clang formattslil clingman
Previously the code base assumed that there was a single, global game state which was the implicit target of all actions taken. Looking ahead at architectural improvements, this has now been (almost entirely) made explicit and functions take tak_state_p where necessary (and also where unnecessary). Two important fixes to actions.c were made: - Previously when generating the possible stack moves, stack height overflows (> 15) were not taken into account and this resulted in the tree search corrupting the board state. Now action search does not list all legal actions, rather the subset of these encodeable by the implementation. - The check for crushing on a stack move was incorrect (too strict), and this resulted in many legitimate moves being igonored. Finally, in other changes, weights have also been improved by training all games instead of some subset for chosen players, and clang-format was run on the codebase.
42 hoursnew neural network arch (faster + better) & minor changes + fixestslil clingman
Gone is the convolutional neural network, for it turns out not only is it more difficult to train, but all of the extra information about board layers didn't make much of a difference at this size. So cnn1986 has been replaced by nn1986, a standard, two-layer, dense nn configured as a binary classifier and (mis)used in that capacity. Note: total number of parameters is unchanged. HARK: this new nn exposes a bug somewhere in ctak. Run ctlm with self-play to see the completely borked board state at the end.
42 hoursWelcome geminict!tslil clingman
This is a special interface to negamax_cnn1986 which is designed to generate output for use in a CGI tak interface to be used over gemini. Also in this commit is a reformating of the various source files to use the traditional tab width of 8 spaces.
42 hoursMore towel wringing: re-implemented check_road_colourtslil clingman
Previously check_win would call check_road_colour once for each road colour, and check_road_colour would call a depth-first search (DFS) for each of the two axes. This meant that we were doing (up to) *four* depth-first searches for each call of check_win. I have replaced both axial DFS with the world's worst TM implementation of a connected component generation algorithm backed by the least guaranteed disjoint set data structure. Essentially doing anything about union find correctly is slower than just ... not doing it. Although we lose the asymptotic complexity, in practice we're doing this millions of times per turn, for a fixed board size and that's what matters. All in all, it appears that i've managed to shave about 69ns off check_win, per call -- nice! This amounts to 50ms or so saved at depth 5 per engine move, in one of my test games. Unfortunately nearly 99% of the time is still taken by evaluating the convolutional neural network. It's slow.
42 hoursTrying to make things fastertslil
I tried the following, but they all made things worse: - moving away from the singly-linked (tail tracking) list for actions by: + using an array zipper for a deque + using an array to poorly hold a floating deque - caching the results of generating move lists in the transposition table and then + copying the resulting list/zip/deque instead of generating it + applying the move-to-front without copying, but this made the search order worse. Presumably in this case shallower nodes were messing up the search tree with garbage moves? I think some of this is not supposed to happen, but i have just the right combination of poor evaluation function and naively ordered and cheap move generation that i'm in a local minimum here.
42 hoursFix copyright notice in files, and small preemptive optimisationtslil clingman
Eventually there'll be a more complicated data generation step than the one we're presently using, so having it in-lined in the loop is wasteful. Ideally also this would be update per ply and we could avoid recalculating it entirely for every query -- though it's probably ``fast enough'' for now. Also, caching is WIP.
42 hoursRename ct_k -> ct, IANAL but ...tslil
42 hoursTEI interface working!tslil clingman
42 hoursTried some naive iterative deepening. Work on TEI interface nexttslil clingman
If TEI is implemented, then i could make use of Morten's racetrack (https://github.com/MortenLohne/racetrack) and develop a quantitative measure of the bot's performance. This is the current priority.
42 hoursJust some #weightgoals ;)tslil clingman
It turns out that while i was training on a 0/1 classification problem, i was using 2*eval - 1. Training using this function instead, and on bot-dominated game choices (chosen_player in extract.sh) seems to have given a better evaluation function. At the least, Morten's swindle doesn't work anymore.
42 hoursFairly important bug fixes to lcdlib, LCD now echoes input!tslil clingman
Input polling without line-buffering is done using ncurses, so the buildroot configuration had to change accordingly to include that library. The Makefile changed to accommodate stand-alone building of ct1986 and to include -lcurses where appropriate. There were also some typos about copying ct1986 and ctaklm to the correct directories.
42 hoursTry to squeeze out a little more performancetslil clingman
``Common wisdom'' dictates that placements are often better than stack moves, so we bias the generated move list in this fashion. Seems to be a little faster.
42 hoursPurged uninteresting statisticstslil
42 hoursAdded license information!tslil clingman
42 hoursPrint progress before recursion rather than aftertslil clingman
42 hoursTypotslil clingman
42 hoursRemoved treap in favour of linked-list chained hash tabletslil clingman
42 hoursIt would appear that any function call whatsoever is slower :/tslil clingman
For now we'll stay with directly recomputing it at each non-terminal node
42 hoursI don't have the presence of mind to debug this right nowtslil clingman
42 hoursSomehting along these lines, i'm tiredtslil clingman
42 hoursSmall changes to build ct1986tslil clingman
42 hoursStoring best moves!tslil clingman
42 hoursThis matches alpha-beta!tslil clingman
42 hoursStable negamax-alpha-beta fail-softtslil clingman
42 hoursOnce again, adding TT changes the outcometslil clingman
42 hoursActually it seems before i was mis-countingtslil clingman
42 hoursThere is still a bug, it doesn't appear to be checking enoughtslil clingman
42 hoursStripping debug stufftslil clingman
42 hoursIt was a silly typo! Hoorah!tslil clingman
42 hoursStill bugs...tslil clingman
42 hoursStill some bugs, standing stone becomes flat at depth4 self-play??tslil clingman
42 hoursThere's still something wrongtslil clingman
42 hoursThis looks better to me and confirms scribbles on papertslil clingman
42 hoursI don't understand why action_list.c:306 != 259tslil clingman
42 hoursThere are still some bugs in the undo almost surely...tslil clingman
42 hoursLots of bugfixes, mostly uint vs int. Still weirdness in gametslil clingman
42 hoursThis is the basic idea, there's ≥ 1 bug (generates illegals...)tslil clingman
42 hoursWorking on action lists to refactortslil clingman
42 hoursStill tryingtslil clingman
|
42 hoursHash collisionstslil clingman
42 hoursFixed some important bugstslil clingman
42 hoursAttempting Zobrist hashingtslil clingman
42 hoursNot correct usage, need to try hashingtslil clingman
42 hoursno speed increase ...tslil clingman
42 hoursTrying caching to speed things uptslil clingman
42 hoursThanks compile warnings!tslil clingman
42 hoursStarting work on LCD librarytslil clingman
42 hoursFinally (?) fixed winning avoidancetslil clingman
42 hoursRenamingtslil clingman
42 hoursMany fixes, i think this is actually correcttslil clingman