Thoughts on learning about AI Safety

I think what I find hard about learning AI safety (and maybe this is even more applicable because of my background in theory) is that there isn’t yet an established theory of AI safety. Most of what I’m reading seems like a selection of research perspectives. And for the part that is more rigorous, the deep learning fundamentals, even there not all design choices are sufficiently motivated by theory but rather empirical results. It is exciting to enter this and hopefully unravel the mess, and it feels possible in the age of LLMs being superhuman personal tutors, but it is also often frustrating. I will try, in an effort to stay sane, to maintain a balance between studying deep learning / probability / statistics foundations and reading various people’s perspectives on approaches to AI safety.