关注
Vladimir Mikulik
Vladimir Mikulik
Anthropic
在 anthropic.com 的电子邮件经过验证
标题
引用次数
引用次数
年份
Gemini: a family of highly capable multimodal models
G Team, R Anil, S Borgeaud, Y Wu, JB Alayrac, J Yu, R Soricut, ...
arXiv preprint arXiv:2312.11805, 2023
13682023
Inferring the effectiveness of government interventions against COVID-19
JM Brauner, S Mindermann, M Sharma, D Johnston, J Salvatier, ...
Science 371 (6531), eabd9338, 2021
10212021
Scaling language models: Methods, analysis & insights from training gopher
JW Rae, S Borgeaud, T Cai, K Millican, J Hoffmann, F Song, J Aslanides, ...
arXiv preprint arXiv:2112.11446, 2021
9242021
Teaching language models to support answers with verified quotes
J Menick, M Trebacz, V Mikulik, J Aslanides, F Song, M Chadwick, ...
arXiv preprint arXiv:2203.11147, 2022
1752022
Alignment of language agents
Z Kenton, T Everitt, L Weidinger, I Gabriel, V Mikulik, G Irving
arXiv preprint arXiv:2103.14659, 2021
1422021
Risks from learned optimization in advanced machine learning systems
E Hubinger, C van Merwijk, V Mikulik, J Skalse, S Garrabrant
arXiv preprint arXiv:1906.01820, 2019
1272019
The DeepMind JAX Ecosystem, 2020
I Babuschkin, K Baumli, A Bell, S Bhupatiraju, J Bruce, P Buchlovsky, ...
URL http://github. com/deepmind 18, 2010
992010
Specification gaming: the flip side of AI ingenuity
V Krakovna, J Uesato, V Mikulik, M Rahtz, T Everitt, R Kumar, Z Kenton, ...
DeepMind Blog 3, 2020
962020
The effectiveness and perceived burden of nonpharmaceutical interventions against COVID-19 transmission: a modelling study with 41 countries
JM Brauner, S Mindermann, M Sharma, AB Stephenson, T Gavenčiak, ...
MedRxiv, 2020.05. 28.20116129, 2020
842020
The DeepMind JAX Ecosystem
I Babuschkin, K Baumli, A Bell, S Bhupatiraju, J Bruce, P Buchlovsky, ...
URL http://github. com/deepmind 24, 25, 2020
622020
Tracr: Compiled transformers as a laboratory for interpretability
D Lindner, J Kramár, S Farquhar, M Rahtz, T McGrath, V Mikulik
Advances in Neural Information Processing Systems 36, 2024
452024
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla
T Lieberum, M Rahtz, J Kramár, N Nanda, G Irving, R Shah, V Mikulik
arXiv preprint arXiv:2307.09458, 2023
432023
Meta-trained agents implement bayes-optimal agents
V Mikulik, G Delétang, T McGrath, T Genewein, M Martic, S Legg, ...
Advances in neural information processing systems 33, 18691-18703, 2020
412020
The hydra effect: Emergent self-repair in language model computations
T McGrath, M Rahtz, J Kramar, V Mikulik, S Legg
arXiv preprint arXiv:2307.15771, 2023
352023
Neural networks are a priori biased towards boolean functions with low entropy
C Mingard, J Skalse, G Valle-Pérez, D Martínez-Rubio, V Mikulik, ...
arXiv preprint arXiv:1909.11522, 2019
282019
Algorithms for causal reasoning in probability trees
T Genewein, T McGrath, G Delétang, V Mikulik, M Martic, S Legg, ...
arXiv preprint arXiv:2010.12237, 2020
202020
Causal analysis of agent behavior for ai safety
G Déletang, J Grau-Moya, M Martic, T Genewein, T McGrath, V Mikulik, ...
arXiv preprint arXiv:2103.03938, 2021
102021
Challenges with unsupervised LLM knowledge discovery
S Farquhar, V Varma, Z Kenton, J Gasteiger, V Mikulik, R Shah
arXiv preprint arXiv:2312.10029, 2023
92023
系统目前无法执行此操作,请稍后再试。
文章 1–18