卡尔纳普为技术目的构建框架时,聚焦于框架的选择与建构。如布劳顿(G. L. Broughton)所言,卡尔纳普关注如何“建构系统并制定规则”。[10]用卡尔纳普的话说,要谈论“新型实体”,必须“引入受新规则支配的新说话方式系统”[11];相关语义规则一旦明确建立,便确定了一类卡尔纳普式指称词项,从而肯定这些词项所指实体的普遍存在。与之相对,托马森的容易论证运作于以自然语言为基础的框架中,该框架被语言共同体所接受。框架本身及其规则既非被设计也非被刻意引入——它们已然存在。她的目标是通过分析自然语言的实际使用来推导本体论结论。因此,与卡尔纳普不同,托马森不涉及框架设计或建构问题。
托马森继承了卡尔纳普的区分,将本体论问题分为“可回答的”与“不可回答的”。[17]然而,她的框架与卡尔纳普存在根本差异。卡尔纳普式框架是形式语义系统,既非自然语言也非其片段,而是仅包含显式定义语言表达式的规整系统;而托马森的自然语言框架则直接采用日常语言中既存的术语。这一区别引入了额外复杂性:卡尔纳普式框架通过显式规定来定义指称词项,而托马森的方法需考察经验性语言事实。因此,在自然语言框架中,卡尔纳普式指称词项可定义如下:语言表达式e在CF中是C-指称词项(卡尔纳普式指称词项),当且仅当CF中存在语义规则规定e指称某实体,或该规则被采用CF的语言共同体隐式接受。此处关键点在于,“指称”必须在本体论而非纯形式语言意义上理解。霍尔(E. W. Hall)与伯格曼(G. Bergmann)曾误将卡尔纳普的纯粹语义学理解为仅规定符号间关系,忽视其本体论维度。对此,卡尔纳普澄清道:“语义学的任务不是发现事实,而是解释语言。尽管限于此任务,语义学确实且必须涉及超语言实体……”[18]关于如“‘a’指称芝加哥”这类语义规则,他进一步解释:“第一项所指涉的——这是决定性的一点——是第二项中出现的名称所指的实体;该实体不是词‘芝加哥’,而是物理对象芝加哥。”[19]
通过区分不同类型的指称词项,我们现在可以考察这些词项在托马森容易论证中的作用及其潜在影响。这个影响出自两个方面。第一,基于自然语言的框架可能同时包含M-S-指称词项和C-指称词项,容易论证前提b中的任何S-指称词项都可能有两种解释。例如,“the number of Fs”可能指称乔姆斯基句法的D中的数M-S(the numberM-S of Fs),或卡尔纳普式系统中的数C(the numberc of Fs,此处的F是一个分类词,比如“苹果”)。第二,引入指称词项的转换规则之间差异显著——有些能确保b中的名词短语是C-指称词项,有些则不能。这两个因素决定了容易论证可能落入三种情景之一,每种情景的有效性需单独检验。
[1] 托马森将容易论证分为三类:她本人的新卡尔纳普式论证、希弗(S. Schiffer)的“无中生有”(something-from-nothing)论证,以及数学哲学中为数的存在辩护的新弗雷格主义论证。托马森与希弗对这些论证的呈现方式相似。(A. L. Thomasson, Ordinary Objects, Oxford: Oxford University Press, 2007;“Existence Questions”, Philosophical Studies 141 (1), 2008, pp. 63-78;“Fictionalism versus Deflationism”, Mind 122 (488), 2013, pp. 1023-1051; Ontology Made Easy, Oxford: Oxford University Press, 2015; “Why We Should Still Take It Easy”, Mind 126 (503), 2017, pp. 769-779.S. Schiffer, “Language-created Language-independent Entities”, Philosophical Topics 24 (1), 1996, pp. 149-167; The Things We Mean, Oxford: Oxford University Press, 2003.)相比之下,新弗雷格主义者主要关注数学本体论,他们试图通过在二阶逻辑中用休谟原则(而非基本公理V)构建一致系统,来支持算术的逻辑主义与柏拉图主义。(具体论证参见B. Hale & C. Wright, The Reason’s Proper Study: Essays towards a Neo-Fregean Philosophy of Mathematics, Oxford: Oxford University Press, 2001.)
[3] S. Yablo, “Go Figure: A Path Through Fictionalism”, Midwest Studies in Philosophy 25(1), 2001, pp. 72-102; “The Myth of the Seven”, Fictionalism in Metaphysics, ed. by M. E. Kalderonm, Oxford: Clarendon Press, 2005, pp. 88-115; Aboutness, Oxford: Princeton University Press, 2014.
[5] A. L. Thomasson, “Fictionalism versus Deflationism”, pp. 1036-1038; A. L. Thomasson, Ontology Made Easy, p. 172.
[6] M. Plebani, “Fictionalism versus Deflationism: A New Look”, Philosophical Studies 175(2), 2018, p. 306.
[7] A. L. Thomasson, Ordinary Objects, p. 167.
[8] 孔蒂萨将紧缩论者面临的问题构述为一个两难困境:要么容易论证的支持者承认其中存在放大推理,从而直接削弱论证效力;要么他们为第一个前提赋予更多内容,但这会使该前提不再具有无可争议性。(G. Contessa, “It Ain’t Easy: Fictionalism, Deflationism, and Easy Arguments in Ontology”, Mind 125 (499), 2016, pp. 763-773.)
[9] A. L. Thomasson, Ontology Made Easy, p. 44.
[10] G. L. Broughton, “Carnapian Frameworks”, Synthese 199 (1-2), 2021, pp. 4097-4126, Secs. 4-5.
[11] R. Carnap, “Empiricism, Semantics, and Ontology”, Meaning and Necessity, 2nd edition, Chicago: University of Chicago Press, 1950/1956, p. 206.
[12] R. Carnap, Introduction to Semantics, Cambridge: Harvard University Press, 1942, pp, 11-12.
[13] A. L. Thomasson, Ontology Made Easy, p. 270.
[14] T. Sider, Writing the Book of the World, Oxford: Oxford University Press, 2011, p. 69.
[15] R. Carnap, “Empiricism, Semantics, and Ontology”, p. 208.
[16] Ibid., pp. 213-214.
[17] A. L. Thomasson, Ontology Made Easy, Secs. 1.1, 3.1.
[18] R. Carnap, “Hall and Bergmann on Semantics”, Mind 54 (214), 1945, p. 148.
[19] Ibid., p. 152.
[20] A. L. Thomasson, Ontology Made Easy, p. 154.
[21] Cf. N. Chomsky, Knowledge of Language: Its Nature, Origin and Use, New York: Praeger Publishers, 1986; N. Chomsky, Lectures on Government and Binding, 5th edition, Dordrecht: Foris Publications, 1988.
[22] N. Chomsky, Knowledge of Language: Its Nature, Origin and Use, pp. 44-45; N. Chomsky, Lectures on Government and Binding, p. 324; N. Chomsky, “Explaining Language Use”, Philosophical Topics 20 (1), 1992, pp. 222-225; N. Chomsky, The Science of Language: Interviews with James McGilvray, Cambridge: Cambridge University Press, 2012, pp. 28-29.
[23] N. Chomsky, Knowledge of Language: Its Nature, Origin and Use, p. 45; N. Chomsky, Lectures on Government and Binding, p. 324.
[24] A. L. Thomasson, Ontology Made Easy, p. 154.
[25] A. L. Thomasson, Ontology Made Easy, pp. 102-103.
[26]Ibid., pp. 221-229.
[27] 假定所有语义规则都表达成条件式。
[28] A. L. Thomasson, Ordinary Objects, Chap.2; A. L. Thomasson, Ontology Made Easy, pp. 222-226, 263-266.
[29] J. Collins, “The Semantics and Ontology of the Average American”, Journal of Semantics 34 (3), 2017, p. 390.
[30] 这个例子参见B. Hale, Abstract Objects, New York: Blackwell, 1987, p. 24.
[31] T. Button, “Deflationary Metaphysics and Ordinary Language”, Synthese 197 (1), 2020, p. 37.
[32] S. Schiffer, The Things We Mean, p. 53.
[33] T. Hofweber, Ontology and the Ambitions of Metaphysics, Oxford: Oxford University Press, 2016, p. 139.
[34] 托马森在此并未直接援引语义规则,而是说如果真有合适的共应用条件被满足,则可接受相应实体的存在。(A. L. Thomasson, Ontology Made Easy, p. 265.)在她看来,共应用条件的满足蕴涵语义规则存在。但正如本文所论证的,这个蕴涵关系是可疑的。
[35] A. L. Thomasson, “What Do Easy Inferences Get Us”, Philosophy and Phenomenological Research 102 (3), 2021, p. 738.
[36] T. Hofweber, Ontology and the Ambitions of Metaphysics, p. 25.
[37] 关于霍夫韦伯对涉及数的容易论证的阐释,参见T. Hofweber, “Innocent Statements and Their Metaphysically Loaded Counterparts”, Philosophers’ Imprint 7(1), 2007, Secs. 3-5, 6.1; T. Hofweber, Ontology and the Ambitions of Metaphysics, pp. 40-44, Chap. 5; T. Hofweber, “Thomasson on Easy Arguments”, M. Garcia-Godinez (ed.), Thomasson on Ontology, London: Palgrave Macmillan, 2023, pp. 47-49. 霍夫韦伯对涉及性质的容易论证的阐释则完全类似。
[38] Arthur Schipper, “Singular Terms and Ontological Seriousness”, Journal of the American Philosophical Association 9 (3), 2023, pp. 574-595.
[39] J. Heil, From an Ontological Point of View, Oxford: Oxford University Press, 2003.
[40] Arthur Schipper, “Aboutness and Ontology: A Modest Approach to Truthmakers”, Philosophical Studies 177 (2), 2020, pp. 513, 527-528; Arthur Schipper, “Singular Terms and Ontological Seriousness”, pp. 576, 579, 592-593.
在一个数学证明中,数学家们选择自己的证明方法的依据是什么?例如,在处理一个几何学的问题时,我们为什么要引入函数的方法?德国数学家希尔伯特(D.Hilbert)在其著作《几何基础》中将上述问题归结为数学证明中“方法的纯粹性”(purity of method)问题。在《几何基础》一书中,希尔伯特创造性地对几何学进行了公理化处理。通过系统地构建几何学的公理体系,希尔伯特为欧氏几何和非欧几何提供了一个严格的逻辑基础,并验证了各公理的独立性、完备性及一致性,为20世纪的数学研究(包括物理学的公理化)奠定了基础。(参见Gray,2007)“虽然哥德尔不完全性定理宣告了纲领的失败,但其核心的有穷证明论和公理化方法影响深远,纲领也以各种改进的方案继续发展。”(康孝军,第44页)其中,“纯粹性”问题源于希尔伯特对几何证明的完整性和一致性的思考,即在证明一些几何定理时,是否仅能使用与定理内容直接相关的概念和工具。(参见Giovannini)特别是,希尔伯特希望避免引入不必要的“外部”概念,以保持证明方法的纯粹性。(参见Venturi)希尔伯特认为:
这个原则,即应当在各处阐明证明可能性的原则,也与“证明方法的纯粹性”这一要求密切相关,这一要求在近代被一些数学家强调过。这一要求其实无非是对所遵循的基本原则的一种主观理解。实际上,前面的几何研究总体上试图揭示,证明一个初等几何真理所需要的公理、假设或工具是什么,最终应由具体情况来决定,从所处的立场出发,哪种证明方法是值得优先选择的。(Hallett and Majer,p.523)
哈雷特(M.Hallett)将这里的纯粹性阐释为证明方法与证明主题是否相适应的问题,即“在数论的证明中使用复变函数理论(正如黎曼和狄利克雷的工作所普及的那样),或者在分析学或点集理论的证明中使用超限数,这种做法是否‘合适’?是否应当‘正确’或‘更好’地采用综合几何而非解析几何?希尔伯特对此问题的回应是,没有哪种方法是‘正确’的,如果有特定的目的,每种方法都是合适的,而‘正确’的方法是同时接受这两种发展方向”(Hallett,p.199)。表面上看,希尔伯特的回应似乎表达了一种在数学方法论上的实用主义立场,即我们对证明目标的理解决定了我们选择证明方法的标准。但哈雷特提醒我们,希尔伯特对“纯粹性”问题的处理并非要寻求一个选择证明方法的“标准”,而是要探究“数学知识的来源”。(参见同上,p.199)换言之,“证明方法的纯粹性”问题揭示了证明方法同通过这种方法所获得的数学知识的合法性之间的关系。例如,在《几何基础》中,希尔伯特一方面认为欧几里得对平行公理(第五公理)的描述过于复杂且缺乏与之相关的独立性证明;另一方面认为欧几里得的公理系统缺乏逻辑上的严谨性,即欧几里得的公理系统中有许多隐含的假设和未被明示的公理,这导致体系内的推理有时依赖于直观理解而非严格的逻辑。基于上述原因,希尔伯特认为欧几里得的公理系统是不完善的。而《几何基础》一书的目标则是通过建立一套完备、一致且独立的公理体系,即这些公理应该足以推导出欧几里得几何的所有已知定理,既不应相互矛盾,且任何公理都不应从其他公理中推导出来(参见Dillon,p.97),来为几何学的知识寻找合法性(基础)。区别于之前的数学家,希尔伯特并不打算在公理系统之外为几何学寻找基础,而是试图在公理之间的逻辑关系中寻找几何知识的合法性根据。(参见Eder and Schiemer)具体来说,希尔伯特通过引入一套精确的公理系统来定义几何学的基本概念(如点、直线、平面等)和关系(如包含关系、顺序关系、平行性等),从而使几何学的基础不再依赖于直观,而是完全基于逻辑推导。换而言之,通过借助严格的逻辑和符号系统,希尔伯特试图将几何学形式化,并以此削弱几何学对直观图形和空间经验的依赖,从而使几何学成为一门形式科学。简言之,“《几何基础》重构了欧氏几何的经典框架,对基本的概念与命题进行了全新的表述,但使用了一种高度形式化的语言。希尔伯特认为,几何命题的数学意义和几何理论的严格性最终都体现在一种公理化架构中,而这种纯粹的结构或关系形式应当用尽可能不带有特定语义的词汇来刻画,因此对传统的欧氏公理系统本身的进一步严密化和抽象化不仅是公理化思想的必然结果,而且只有这样才能表明真正的‘数学性’”(钱立卿,第153—154页)。
为了进一步化简计算,拉普拉斯又引入了一个新的变量,用这个变量来重新表示原积分中的各项,并计算其导数以调整积分的结构。通过变量替换重新构造积分的形式,拉普拉斯简化了原积分中的复杂幂次和指数项,使得结果可以标准化,甚至与已知的特殊函数(如Gamma函数)相关联,从而降低了处理该积分的难度。(参见同上)然而,泊松却认为,使用复数作为工具来求解实数积分问题是一种“间接”的方法,其与问题的本质不符。基于这一立场,泊松完全立足于实数分析的方法来解决这一问题。他通过复杂的计算和冗长的步骤,最终推导出与拉普拉斯相同的结果。(参见Poisson)佩尔(B.Pel)则认为,虽然泊松所要求的“直接性”(directness)确实是现代数学哲学语境中的“纯粹性”(purity),但绝非严格意义上的纯粹性。(参见Pel)泊松认为,尽管我们能够通过拉普拉斯的方法得到正确答案,但它不是“直接”的,因为拉普拉斯在证明中引入了复数,而这超出了实数分析的范围。这种要求证明的每个步骤都必须与证明问题的主题相一致的纯粹性是一种“严格的纯粹性”,或德特勒夫森(M.Detlefsen)和阿拉纳所说的“主题的纯粹性”(topical purity)。(参见Poisson;Detlefsen and Arana)。在佩尔看来,如果把“主题的纯粹性”当作标准,那么即使是泊松的证明也不满足这一纯粹性的要求。因为泊松在他的证明中使用了双重积分,但理解原始问题却不需要定义双重积分。(参见Pel)基于这一分析,佩尔认为泊松的直接性更接近于卡勒(R.Kahle)和普尔奇尼(G.Pulcini)所说的“操作的纯粹性”(operational purity)②。(参见Pel;Kahle and Pulcini)操作的纯粹性关注数学证明中方法的层次性和匹配性,而不仅仅是逻辑上的正确性。卡勒和普尔奇尼试图通过一个方法所应用的对象在本体论层面的“高低”来判断该方法是否纯粹。为了说明这一纯粹性标准,卡勒和普尔奇尼首先区分了一个证明中的“操作内容”(operational content)及其“本体论”(ontology)。(参见Kahle and Pulcini,p.134)“操作内容”指定理或证明中涉及的数学操作或被运用到的方法的集合,例如加法、乘法、除法、积分等。此外,每个操作都依赖一个对该操作封闭的数域。所谓“封闭”则指对数域中的任意元素进行某种运算后得到的结果仍然属于这个数域。例如,自然数N对加法和乘法是封闭的,但对减法和除法是不封闭的,因为两个自然数相减或相除可能得到一个不属于“自然数”这个数域的数,如2-3=-1或3÷2=1.5。“本体论”则指“操作内容”所依赖的最小封闭域,例如加法的最小封闭域是“自然数”,而“整数”对减法封闭,“有理数”则对除法封闭等。此外,不同方法所属的层次也不同。例如,相较于加法,除法被视为一种更高层次的方法,因为自然数在加法运算下是封闭的,也就是说,自然数的加法运算不会产生超出自然数范围的结果;而在除法运算中,为了保证结果的封闭性,必须引入一个更广泛的数域,即包含所有自然数的有理数集合。(参见Pel)因此,如果一个证明的本体论是其所出发的定理的本体论的子集,那么这个证明就被认为是在方法上纯粹的。
潘萨论证的关键在于找到(i)与(ii)的共同特征,并以此得出一个检验(ii)中证明的标准。对此,潘萨认为可以找到以下五个特征:(1)它们都是由一系列有序的、可重复的人类活动(acts)所构成,且以某种按照特定时空顺序排列的符号配置(configuration of signs)为基础;(2)它们都通过既有的符号配置获得了某些证据(evidence),且这种证据被认为是进一步生成新的符号配置的合理依据(warrant);(3)它们的有效性并不完全依赖于普遍的逻辑规则,而是依赖于证明中特定的上下文和情境;(4)每个具体的数学证明都必须在特定的理论框架和操作规则下进行,且这些规则并非普遍适用于所有数学证明;(5)它们都是一种交流活动(activity of communication),即主体间借助不同的符号(以听觉符号和图形符号为主的)系统所进行的交流过程——笔者将这一特征称之为“可互动性”。在上述特征中,一个证明(无论该证明是否形式化)是否能被理解和承认,很大程度上取决于它是否具有可互动性。原因在于,数学证明不仅是一个个体内在的逻辑演绎过程,而且是一个交流和共享的过程。这一过程涉及数学家之间的互动、解释、辩论以及共识的达成。那么,进一步的问题在于,我们应该如何检验一个证明是否具有可互动性?事实上,这涉及其他数学家对某个证明的理解过程。哈马米(Y.Hamami)和莫里斯(R.L.Morris)认为,理解一个数学证明的过程就是理性地重构其“基础计划”(underlying plan)的过程。(参见Hamami and Morris)这一观点为重新界定证明的“可理解性”提供了理论基础。理解数学证明不仅意味着逐步检验每一条推理是否有效,更关键的是追踪并把握整个证明的结构性意图,即解释每一步推理在整体论证中所承担的角色与功能。这一过程本质上是具有层次性和动态性的:既涉及形式步骤的完备性,也取决于理解者是否能够识别出一条连贯且有目的的推理轨迹。
[3]d’Alembert,J.L.R.,1751,”Application de l’algebre ou de l’analyse à la géométrie”,in D.Diderot and J.L.R.d’Alembert(eds.),Encyclopédie ou dictionnaire raisonné des sciences,des arts et des métiers,vol.1,Paris:Briasson,David,Le Breton,and Durand.
[4]Arana,A.,2014,”Purity in Arithmetic:Some Formal and Informal Issues”,in Formalism and beyond.On the Nature of Mathematical Discourse.
2017,”On the Alleged Simplicity of Impure Proof”,in Simplicity:Ideals of Practice in Mathematics and the Arts.
2024,Elements of Purity.Elements in the Philosophy of Mathematics,Cambridge:Cambridge University Press.
[6]Baldwin,J.T.,2013,”Formalization,Primitive Concepts,and Purity”,in The Review of Symbolic Logic 6(1).
[7]Compton,T.,1990,”What Are the TOPNOI in Philebus 51c”,in The Classical Quarterly 40(2).
[8]Detlefsen,M.,1990,”On an Alleged Refutation of Hilbert’s Program Using Gödel’s First Incompleteness Theorem”,in Journal of Philosophical Logic 19(4).
1996,”Philosophy of Mathematics in the Twentieth Century”,in Philosophy of Science,Logic,and Mathematics in the Twentieth Century,S.G.Shanker(ed.),vol.9,Routledge History of Philosophy,London and NY:Routledge.
[9]Detlefsen,M.and Arana,A.,2011,”Purity of Methods”,in Philosophers 11(2).
[10]Dillon,M.I.,2018,”Hilbert’s Grundlagen”,in Geometry Through History:Euclidean,Hyperbolic,and Projective Geometries,Springer International Publishing AG.
[11]Eder,G.and Schiemer,G.,2017,”Hilbert,Duality,and the Geometrical Roots of Model Theory”,in The Review of Symbolic Logic 11.
[12]Fu,L.,2025,”Beauty Loads to Truth:Aesthetic Induction on Consistency”,in Synthese 206(44).
[13]Giovannini,E.N.,2021,”David Hilbert and the Foundations of the Theory of Plane Area”,in Archive for History of Exact Sciences 75(6).
[14]Gray,J.,2007,Worlds Out of Nothing:A Course in the History of Geometry in the 19th Century,vol.193,London:Springer.
[15]Hallett,M.,2008,”Reflections on the Purity of Method in Hilbert’s Grundlagen der Geometric”,in P.Mancosu(ed.),The Philosophy of Mathematical Practice,Oxford:OUP Oxford.
[16]Hallett,M.and Majer,U.(eds.),2004,David Hilbert’s Lectures on the Foundations of Geometry 1891-1902,vol.1,New York:Springer Science & Business Media.
[17]Hamami,Y.and Morris,R.L.,2024,”Understanding in Mathematics:The Case of Mathematical Proofs”,in Noûs 58.
[18]Hayn-Leichsenring,G.U.,Vartanian,O.,and Chatterjee,A.,2022,”The Role of Expertise in the Aesthetic Evaluation of Mathematical Equations”,in Psychological Research 86(5).
[19]Kahle,R.and Pulcini,G.,2018,”Towards an Operational View of Purity”,in P.Arazim and T.Pávička(eds.),The Logica Yearbook 2017,London:College Publications.
[20]Laplace,P.-S.,1809,”Mémoire sur divers points d’analyse”,in Oeuvres Complètes,vol.14.
[21]Lemhoff,R.,2017,”Remarks on Simple Proofs”,in R.Kossak and P.Ording(eds.),Simplicity:Ideals of Practice in Mathematics and the Arts,Springer International Publishing AG.
[22]Lennon,T.M.,2005,”Percorsi anticartesiani nelle lettere a Pierre-Daniel Huet(review)”,in Renaissance Quarterly 58.
[23]Luchins,A.S.and Luchins,E.H.,1990,”The Einstein-Wertheimer Correspondence on Geometric Proofs and Mathematical Puzzles”,in The Mathematical Intelligencer 12.
[24]Newton,I.,1967,”Universal Arithmetick”,in D.T.Whiteside(ed.),The Mathematical Works of Isaac Newton,vol.Ⅱ,NY:Johnson Reprint Corporation.
Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, Katsushi Ikeuchi, Hoi Vo, Li Fei-Fei1, Jianfeng Gao
Figure 1: Overview of an Agent AI system that can perceive and act in different domains and applications. Agent AI is emerging as a promising avenue toward Artificial General Intelligence (AGI). Agent AI training has demonstrated the capacity for multi-modal understanding in the physical world. It provides a framework for reality-agnostic training by leveraging generative AI alongside multiple independent data sources. Large foundation models trained for agent and action-related tasks can be applied to physical and virtual worlds when trained on cross-reality data. We present the general overview of an Agent AI system that can perceive and act in many different domains and applications, possibly serving as a route towards AGI using an agent paradigm.
ABSTRACT
Multi-modal AI systems will likely become a ubiquitous presence in our everyday lives. A promising approach to making these systems more interactive is to embody them as agents within physical and virtual environments. At present, systems leverage existing foundation models as the basic building blocks for the creation of embodied agents. Embedding agents within such environments facilitates the ability of models to process and interpret visual and contextual data, which is critical for the creation of more sophisticated and context-aware AI systems. For example, a system that can perceive user actions, human behavior, environmental objects, audio expressions, and the collective sentiment of a scene can be used to inform and direct agent responses within the given environment. To accelerate research on agent-based multimodal intelligence, we define “Agent AI” as a class of interactive systems that can perceive visual stimuli, language inputs, and other environmentally-grounded data, and can produce meaningful embodied actions. In particular, we explore systems that aim to improve agents based on next-embodied action prediction by incorporating external knowledge, multi-sensory inputs, and human feedback. We argue that by developing agentic AI systems in grounded environments, one can also mitigate the hallucinations of large foundation models and their tendency to generate environmentally incorrect outputs. The emerging field of Agent AI subsumes the broader embodied and agentic aspects of multimodal interactions. Beyond agents acting and interacting in the physical world, we envision a future where people can easily create any virtual reality or simulated scene and interact with agents embodied within the virtual environment.
Contents 1 Introduction 1.1 Motivation 1.2 Background 1.3 Overview 2 Agent AI Integration 2.1 Infinite AI agent 2.2 Agent AI with Large Foundation Models 2.2.1 Hallucinations 2.2.2 Biases and Inclusivity 2.2.3 Data Privacy and Usage 2.2.4 Interpret ability and Explain ability 2.2.5 Inference Augmentation 2.2.6 Regulation 2.3 Agent AI for Emergent Abilities 3 Agent AI Paradigm 3.1 LLMs and VLMs 3.2 Agent Transformer Definition 3.3 Agent Transformer Creation 4 Agent AI Learning 4.1 Strategy and Mechanism 4.1.1 Reinforcement Learning(RL) 4.1.2 Imitation Learning(IL) 4.1.3 Traditional RGB 4.1.4 In-context Learning 4.1.5 Optimization in the Agent System 4.2 Agent Systems(zero-shot and few-shot level) 4.2.1 Agent Modules 4.2.2 Agent Infrastructure 4.3 Agentic Foundation Models(pretraining and fine tune level) 5 Agent AI Categorization 5.1 Generalist Agent Areas 5.2 Embodied Agents 5.2.1 Action Agents 5.2.2 Interactive Agents 5.3 Simulation and Environments Agents 5.4 Generative Agents 5.4.1 AR/VR/mixed-reality Agents 5.5 Knowledge and Logical Inference Agents 5.5.1 Knowledge Agent 5.5.2 Logic Agents 5.5.3 Agents for Emotional Reasoning 5.5.4 Neuro-Symbolic Agents 5.6 LLMs and VLMs Agent 6 Agent AI Application Tasks 6.1 Agents for Gaming 6.1.1 NPC Behavior 6.1.2 Human-NPC Interaction 6.1.3 Agent-based Analysis of Gaming 6.1.4 Scene Synthesis for Gaming 6.1.5 Experiments and Results 6.2 Robotics 6.2.1 LLM/VLM Agent for Robotics. 6.2.2 Experiments and Results 6.3 Healthcare 6.3.1 Current Healthcare Capabilities 6.4 Multimodal Agents 6.4.1 Image-Language Understanding and Generation 6.4.2 Video and Language Understanding and Generation 6.4.3 Experiments and Results 6.5 Video-language Experiments 6.6 Agent for NLP 6.6.1 LLM agent 6.6.2 General LLM agent 6.6.3 Instruction-following LLM agents 6.6.4 Experiment sand Results 7 Agent AI Across Modalities, Domains and Realities 7.1 Agents for Cross-modal Understanding 7.2 Agents for Cross-domain Understanding 7.3 Interactive agent for cross-modality and cross-reality 7.4 Sim to Real Transfer 8 Continuous and Self-improvement for Agent AI 8.1 Human-based Interaction Data 8.2 Foundation Model Generated Data 9 Agent Dataset and Leaderboard 9.1 “CuisineWorld” Dataset for Multi-agent Gaming 9.1.1 Benchmark 9.1.2 Task 9.1.3 Metrics and Judging 9.1.4 Evaluation 9.2 Audio-Video-Language Pre-training Dataset 10 Broader Impact Statement 11 Ethical Considerations 12 Diversity Statement
Historically, AI systems were defined at the 1956 Dartmouth Conference as artificial life forms that could collect information from the environment and interact with it in useful ways. Motivated by this definition, Minsky’s MIT group built in 1970 a robotics system, called the “Copy Demo,” that observed “blocks world” scenes and successfully reconstructed the observed polyhedral block structures. The system, which comprised observation, planning, and manipulation modules, revealed that each of these subproblems is highly challenging and further research was necessary. The AI field fragmented into specialized subfields that have largely independently made great progress in tackling these and other problems, but over-reductionism has blurred the overarching goals of AI research.
To advance beyond the status quo, it is necessary to return to AI fundamentals motivated by Aristotelian Holism. Fortunately, the recent revolution in Large Language Models (LLMs) and Visual Language Models (VLMs) has made it possible to create novel AI agents consistent with the holistic ideal. Seizing upon this opportunity, this article explores models that integrate language proficiency, visual cognition, context memory, intuitive reasoning, and adaptability. It explores the potential completion of this holistic synthesis using LLMs and VLMs. In our exploration, we also revisit system design based on Aristotle’s Final Cause, the teleological “why the system exists”, which may have been overlooked in previous rounds of AI development.
With the advent of powerful pretrained LLMs and VLMs, a renaissance in natural language processing and computer vision has been catalyzed. LLMs now demonstrate an impressive ability to decipher the nuances of real-world linguistic data, often achieving abilities that parallel or even surpass human expertise (OpenAI, 2023). Recently, researchers have shown that LLMs may be extended to act as agents within various environments, performing intricate actions and tasks when paired with domain-specific knowledge and modules (Xi et al., 2023). These scenarios, characterized by complex reasoning, understanding of the agent’s role and its environment, along with multi-step planning, test the agent’s ability to make highly nuanced and intricate decisions within its environmental constraints (Wu et al., 2023; Meta Fundamental AI Research (FAIR) Diplomacy Team et al., 2022).
Building upon these initial efforts, the AI community is on the cusp of a significant paradigm shift, transitioning from creating AI models for passive, structured tasks to models capable of assuming dynamic, agentic roles in diverse and complex environments. In this context, this article investigates the immense potential of using LLMs and VLMs as agents, emphasizing models that have a blend of linguistic proficiency, visual cognition, contextual memory, intuitive reasoning, and adaptability. Leveraging LLMs and VLMs as agents, especially within domains like gaming, robotics, and healthcare, promises not just a rigorous evaluation platform for state-of-the-art AI systems, but also foreshadows the transformative impacts that Agent-centric AI will have across society and industries. When fully harnessed, agentic models can redefine human experiences and elevate operational standards. The potential for sweeping automation ushered in by these models portends monumental shifts in industries and socio-economic dynamics. Such advancements will be intertwined with multifaceted leader-board, not only technical but also ethical, as we will elaborate upon in Section 11. We delve into the overlapping areas of these sub-fields of Agent AI and illustrate their interconnectedness in Fig.1.
1.2 Background
We will now introduce relevant research papers that support the concepts, theoretical background, and modern implementations of Agent AI.
Large Foundation Models: LLMs and VLMs have been driving the effort to develop general intelligent machines (Bubeck et al., 2023; Mirchandani et al., 2023). Although they are trained using large text corpora, their superior problem-solving capacity is not limited to canonical language processing domains. LLMs can potentially tackle complex tasks that were previously presumed to be exclusive to human experts or domain-specific algorithms, ranging from mathematical reasoning (Imani et al., 2023; Wei et al., 2022; Zhu et al., 2022) to answering questions of professional law (Blair-Stanek et al., 2023; Choi et al., 2023; Nay, 2022). Recent research has shown the possibility of using LLMs to generate complex plans for robots and game AI (Liang et al., 2022; Wang et al., 2023a,b; Yao et al., 2023a; Huang et al., 2023a), marking an important milestone for LLMs as general-purpose intelligent agents.
Embodied AI: A number of works leverage LLMs to perform task planning (Huang et al., 2022a; Wang et al., 2023b; Yao et al., 2023a; Li et al., 2023a), specifically the LLMs’ WWW-scale domain knowledge and emergent zero-shot embodied abilities to perform complex task planning and reasoning. Recent robotics research also leverages LLMsto perform task planning (Ahn et al., 2022a; Huang et al., 2022b; Liang et al., 2022) by decomposing natural language instruction into a sequence of subtasks, either in the natural language form or in Python code, then using a low-level controller to execute these subtasks. Additionally, they incorporate environmental feedback to improve task performance (Huang et al., 2022b), (Liang et al., 2022), (Wang et al., 2023a), and (Ikeuchi et al., 2023).
Interactive Learning: AI agents designed for interactive learning operate using a combination of machine learning techniques and user interactions. Initially, the AI agent is trained on a large dataset. This dataset includes various types of information, depending on the intended function of the agent. For instance, an AI designed for language tasks would be trained on a massive corpus of text data. The training involves using machine learning algorithms, which could include deep learning models like neural networks. These training models enable the AI to recognize patterns, make predictions, and generate responses based on the data on which it was trained. The AI agent can also learn from real-time interactions with users. This interactive learning can occur in various ways: 1) Feedback-based learning: The AI adapts its responses based on direct user feedback (Li et al., 2023b; Yu et al., 2023a; Parakh et al., 2023; Zha et al., 2023; Wake et al., 2023a,b,c). For example, if a user corrects the AI’s response, the AI can use this information to improve future responses (Zha et al., 2023; Liu et al., 2023a). 2) Observational Learning: The AI observes user interactions and learns implicitly. For example, if users frequently ask similar questions or interact with the AI in a particular way, the AI might adjust its responses to better suit these patterns. It allows the AI agent to understand and process human language, multi-model setting, interpret the cross reality-context, and generate human-users’ responses. Over time, with more user interactions and feedback, the AI agent’s performance generally continuous improves. This process is often supervised by human operators or developers who ensure that the AI is learning appropriately and not developing biases or incorrect patterns.
1.3 Overview
Multimodal Agent AI (MAA) is a family of systems that generate effective actions in a given environment based on the understanding of multimodal sensory input. With the advent of Large Language Models (LLMs) and Vision Language Models (VLMs), numerous MAA systems have been proposed in fields ranging from basic research to applications. While these research areas are growing rapidly by integrating with the traditional technologies of each domain (e.g., visual question answering and vision-language navigation), they share common interests such as data collection, benchmarking, and ethical perspectives. In this paper, we focus on the some representative research areas of MAA, namely multimodality, gaming (VR/AR/MR), robotics, and healthcare, and we aim to provide comprehensive knowledge on the common concerns discussed in these fields. As a result we expect to learn the fundamentals of MAA and gain insights to further advance their research. Specific learning outcomes include:
•MAA Overview: A deep dive into its principles and roles in contemporary applications, providing researcher with a thorough grasp of its importance and uses. •Methodologies: Detailed examples of how LLMs and VLMs enhance MAAs, illustrated through case studies in gaming, robotics, and healthcare. •Performance Evaluation: Guidance on the assessment of MAAs with relevant datasets, focusing on their effectiveness and generalization. •Ethical Considerations: A discussion on the societal impacts and ethical leader-board of deploying Agent AI, highlighting responsible development practices. •Emerging Trends and Future leader-board: Categorize the latest developments in each domain and discuss the future directions.
Computer-based action and generalist agents (GAs) are useful for many tasks. A GA to become truly valuable to its users, it can natural to interact with, and generalize to a broad range of contexts and modalities. We aims to cultivate a vibrant research ecosystem and create a shared sense of identity and purpose among the Agent AI community. MAA has the potential to be widely applicable across various contexts and modalities, including input from humans. Therefore, we believe this Agent AI area can engage a diverse range of researchers, fostering a dynamic Agent AI community and shared goals. Led by esteemed experts from academia and industry, we expect that this paper will be an interactive and
enriching experience, complete with agent instruction, case studies, tasks sessions, and experiments discussion ensuring a comprehensive and engaging learning experience for all researchers.
This paper aims to provide general and comprehensive knowledge about the current research in the field of Agent AI. To this end, the rest of the paper is organized as follows. Section 2 outlines how Agent AI benefits from integrating with related emerging technologies, particularly large foundation models. Section 3 describes a new paradigm and framework that we propose for training Agent AI. Section 4 provides an overview of the methodologies that are widely used in the training of Agent AI. Section 5 categorizes and discusses various types of agents. Section 6 introduces Agent AI applications in gaming, robotics, and healthcare. Section 7 explores the research community’s efforts to develop a versatile Agent AI, capable of being applied across various modalities, domains, and bridging the sim-to-real gap. Section 8 discusses the potential of Agent AI that not only relies on pre-trained foundation models, but also continuously learns and self-improves by leveraging interactions with the environment and users. Section 9 introduces our new datasets that are designed for the training of multimodal Agent AI. Section 11 discusses the hot topic of the ethics consideration of AI agent, limitations, and societal impact of our paper.
2 Agent AI Integration
Foundation models based on LLMs and VLMs, as proposed in previous research, still exhibit limited performance in the area of embodied AI, particularly in terms of understanding, generating, editing, and interacting within unseen environments or scenarios (Huang et al., 2023a; Zeng et al., 2023). Consequently, these limitations lead to sub-optimal outputs from AI agents. Current agent-centric AI modeling approaches focus on directly accessible and clearly defined data (e.g. text or string representations of the world state) and generally use domain and environment-independent patterns learned from their large-scale pretraining to predict action outputs for each environment (Xi et al., 2023; Wang et al., 2023c; Gong et al., 2023a; Wu et al., 2023). In (Huang et al., 2023a), we investigate the task of knowledge-guided collaborative and interactive scene generation by combining large foundation models, and show promising results that indicate knowledge-grounded LLM agents can improve the performance of 2D and 3D scene understanding, generation, and editing, alongside with other human-agent interactions (Huang et al., 2023a). By integrating an Agent AI framework, large foundation models are able to more deeply understand user input to form a complex and adaptive HCI system. Emergent ability of LLM and VLM works invisible in generative AI, embodied AI, knowledge augmentation for multi-model learning, mix-reality generation, text to vision editing, human interaction for 2D/3D simulation in gaming or robotics tasks. Agent AI recent progress in foundation models present an imminent catalyst for unlocking general intelligence in embodied agents. The large action models, or agent-vision-language models open new possibilities for general-purpose embodied systems such as planning, problem-solving and learning in complex environments. Agent AI test further step in metaverse, and route the early version of AGI.
2.1 Infinite AI agent
AI agents have the capacity to interpret, predict, and respond based on its training and input data. While these capabilities are advanced and continually improving, it’s important to recognize their limitations and the influence of the underlying data they are trained on. AI agent systems generally possess the following abilities: 1) Predictive Modeling: AI agents can predict likely outcomes or suggest next steps based on historical data and trends. For instance, they might predict the continuation of a text, the answer to a question, the next action for a robot, or the resolution of a scenario. 2) Decision Making: In some applications, AI agents can make decisions based on their inferences. Generally, the agent will base their decision on what is most likely to achieve a specified goal. For AI applications like recommendation systems, an agent can decide what products or content to recommend based on its inferences about user preferences. 3) Handling Ambiguity: AI agents can often handle ambiguous input by inferring the most likely interpretation based on context and training. However, their ability to do so is limited by the scope of their training data and algorithms. 4) Continuous Improvement: While some AI agents have the ability to learn from new data and interactions, many large language models do not continuously update their knowledge-base or internal representation after training. Their inferences are usually based solely on the data that was available up to the point of their last training update.
We show augmented interactive agents for multi-modality and cross reality-agnostic integration with an emergence mechanism in Fig. 2. An AI agent requires collecting extensive training data for every new task, which can be costly or impossible for many domains. In this study, we develop an infinite agent that learns to transfer memory information from general foundation models (e.g., GPT-X, DALL-E) to novel domains or scenarios for scene understanding, generation, and interactive editing in physical or virtual worlds.
Figure 2: The multi-model agent AI for 2D/3D embodied generation and editing interaction in cross-reality.
An application of such an infinite agent in robotics is RoboGen (Wang et al., 2023d). In this study, the authors propose a pipeline that autonomously run the cycles of task proposition, environment generation, and skill learning. RoboGen is an effort to transfer the knowledge embedded in large models to robotics.
2.2 Agent AI with Large Foundation Models
Recent studies have indicated that large foundation models play a crucial role in creating data that act as benchmarks for determining the actions of agents within environment-imposed constraints. For example, using foundation models for robotic manipulation (Black et al., 2023; Ko et al., 2023) and navigation (Shah et al., 2023a; Zhou et al., 2023a). To illustrate, Black et al. employed an image-editing model as a high-level planner to generate images of future sub-goals, thereby guiding low-level policies (Black et al., 2023). For robot navigation, Shah et al. proposed a system that employs a LLMtoidentify landmarks from text and a VLM to associate these landmarks with visual inputs, enhancing navigation through natural language instructions (Shah et al., 2023a).
There is also growing interest in the generation of conditioned human motions in response to language and environmental factors. Several AI systems have been proposed to generate motions and actions that are tailored to specific linguistic instructions (Kim et al., 2023; Zhang et al., 2022; Tevet et al., 2022) and to adapt to various 3D scenes (Wang et al., 2022a). This body of research emphasizes the growing capabilities of generative models in enhancing the adaptability and responsiveness of AI agents across diverse scenarios.
2.2.1 Hallucinations
Agents that generate text are often prone to hallucinations, which are instances where the generated text is nonsensical or unfaithful to the provided source content (Raunak et al., 2021; Maynez et al., 2020). Hallucinations can be split into two categories, intrinsic and extrinsic (Ji et al., 2023). Intrinsic hallucinations are hallucinations that are contradictory to the source material, whereas extrinsic hallucinations are when the generated text contains additional information that was not originally included in the source material.
Some promising routes for reducing the rate of hallucination in language generation involve using retrieval-augmented generation (Lewis et al., 2020; Shuster et al., 2021) or other methods for grounding natural language outputs via external knowledge retrieval (Dziri et al., 2021; Peng et al., 2023). Generally, these methods seek to augment language generation by retrieving additional source material and by providing mechanisms to check for contradictions between the generated response and the source material.
Within the context of multi-modal agent systems, VLMs have been shown to hallucinate as well (Zhou et al., 2023b). One common cause of hallucination for vision-based language-generation is due to the over-reliance on co-occurrence of objects and visual cues in the training data (Rohrbach et al., 2018). AI agents that exclusively rely upon pretrained LLMs or VLMs and use limited environment-specific finetuning can be particularly vulnerable to hallucinations since they rely upon the internal knowledge-base of the pretrained models for generating actions and may not accurately understand the dynamics of the world state in which they are deployed.
2.2.2 Biases and Inclusivity
AI agents based on LLMs or LMMs (large multimodal models) have biases due to several factors inherent in their design and training process. When designing these AI agents, we must be mindful of being inclusive and aware of the needs of all end users and stakeholders. In the context of AI agents, inclusivity refers to the measures and principles
employed to ensure that the agent’s responses and interactions are inclusive, respectful, and sensitive to a wide range of users from diverse backgrounds. We list key aspects of agent biases and inclusivity below.
•Training Data: Foundation models are trained on vast amounts of text data collected from the internet, including books, articles, websites, and other text sources. This data often reflects the biases present in human society, and the model can inadvertently learn and reproduce these biases. This includes stereotypes, prejudices, and slanted viewpoints related to race, gender, ethnicity, religion, and other personal attributes. In particular, by training on internet data and often only English text, models implicitly learn the cultural norms of Western, Educated, Industrialized, Rich, and Democratic (WEIRD) societies (Henrich et al., 2010) who have a disproportionately large internet presence. However, it is essential to recognize that datasets created by humans cannot be entirely devoid of bias, since they frequently mirror the societal biases and the predispositions of the individuals who generated and/or compiled the data initially.
•Historical and Cultural Biases: AI models are trained on large datasets sourced from diverse content. Thus, the training data often includes historical texts or materials from various cultures. In particular, training data from historical sources may contain offensive or derogatory language representing a particular society’s cultural norms, attitudes, and prejudices. This can lead to the model perpetuating outdated stereotypes or not fully understanding contemporary cultural shifts and nuances.
•Language and Context Limitations: Language models might struggle with understanding and accurately representing nuances in language, such as sarcasm, humor, or cultural references. This can lead to misinterpretations or biased responses in certain contexts. Furthermore, there are many aspects of spoken language that are not captured by pure text data, leading to a potential disconnect between human understanding of language and how models understand language.
•Policies and Guidelines: AI agents operate under strict policies and guidelines to ensure fairness and inclusivity. For instance, in generating images, there are rules to diversify depictions of people, avoiding stereotypes related to race, gender, and other attributes.
•Overgeneralization: These models tend to generate responses based on patterns seen in the training data. This can lead to overgeneralizations, where the model might produce responses that seem to stereotype or make broad assumptions about certain groups.
•Constant Monitoring and Updating: AI systems are continuously monitored and updated to address any emerging biases or inclusivity issues. Feedback from users and ongoing research in AI ethics play a crucial role in this process.
•Amplification of Dominant Views: Since the training data often includes more content from dominant cultures or groups, the model may be more biased towards these perspectives, potentially underrepresenting or misrepresenting minority viewpoints. •Ethical and Inclusive Design: AI tools should be designed with ethical considerations and inclusivity as core principles. This includes respecting cultural differences, promoting diversity, and ensuring that the AI does not perpetuate harmful stereotypes.
•User Guidelines: Users are also guided on how to interact with AI in a manner that promotes inclusivity and respect. This includes refraining from requests that could lead to biased or inappropriate outputs. Furthermore, it can help mitigate models learning harmful material from user interactions.
Despite these measures, AI agents still exhibit biases. Ongoing efforts in agent AI research and development are focused on further reducing these biases and enhancing the inclusivity and fairness of agent AI systems. Efforts to Mitigate Biases:
•Diverse and Inclusive Training Data: Efforts are made to include a more diverse and inclusive range of sources in the training data.
•Bias Detection and Correction: Ongoing research focuses on detecting and correcting biases in model responses.
•Ethical Guidelines and Policies: Models are often governed by ethical guidelines and policies designed to mitigate biases and ensure respectful and inclusive interactions.
•Diverse Representation: Ensuring that the content generated or the responses provided by the AI agent represent a wide range of human experiences, cultures, ethnicities, and identities. This is particularly relevant in scenarios like image generation or narrative construction.
•Bias Mitigation: Actively working to reduce biases in the AI’s responses. This includes biases related to race, gender, age, disability, sexual orientation, and other personal characteristics. The goal is to provide fair and balanced responses that do not perpetuate stereotypes or prejudices.
•Cultural Sensitivity: The AI is designed to be culturally sensitive, acknowledging and respecting the diversity of cultural norms, practices, and values. This includes understanding and appropriately responding to cultural references and nuances.
•Accessibility: Ensuring that the AI agent is accessible to users with different abilities, including those with disabilities. This can involve incorporating features that make interactions easier for people with visual, auditory, motor, or cognitive impairments.
•Language-based Inclusivity: Providing support for multiple languages and dialects to cater to a global user base, and being sensitive to the nuances and variations within a language (Liu et al., 2023b).
•Ethical and Respectful Interactions: The Agent is programmed to interact ethically and respectfully with all users, avoiding responses that could be deemed offensive, harmful, or disrespectful.
•User Feedback and Adaptation: Incorporating user feedback to continually improve the inclusivity and effectiveness of the AI agent. This includes learning from interactions to better understand and serve a diverse user base.
•Compliance with Inclusivity Guidelines: Adhering to established guidelines and standards for inclusivity in AI agent, which are often set by industry groups, ethical boards, or regulatory bodies.
Despite these efforts, it’s important to be aware of the potential for biases in responses and to interpret them with critical thinking. Continuous improvements in AI agent technology and ethical practices aim to reduce these biases over time. One of the overarching goals for inclusivity in agent AI is to create an agent that is respectful and accessible to all users, regardless of their background or identity.
2.2.3 Data Privacy and Usage
One key ethical consideration of AI agents involves comprehending how these systems handle, store, and potentially retrieve user data. We discuss key aspects below:
Data Collection, Usage and Purpose. When using user data to improve model performance, model developers access the data the AI agent has collected while in production and interacting with users. Some systems allow users to view their data through user accounts or by making a request to the service provider. It is important to recognize what data the AI agent collects during these interactions. This could include text inputs, user usage patterns, personal preferences, and sometimes more sensitive personal information. Users should also understand how the data collected from their interactions is used. If, for some reason, the AI holds incorrect information about a particular person or group, there should be a mechanism for users to help correct this once identified. This is important for both accuracy and to be respectful of all users and groups. Common uses for retrieving and analyzing user data include improving user interaction, personalizing responses, and system optimization. It is extremely important for developers to ensure the data is not used for purposes that users have not consented to, such as unsolicited marketing.
Storage and Security. Developers should know where the user interaction data is stored and what security measures are in place to protect it from unauthorized access or breaches. This includes encryption, secure servers, and data protection protocols. It is extremely important to determine if agent data is shared with third parties and under what conditions. This should be transparent and typically requires user consent.
Data Deletion and Retention. It is also important for users to understand how long user data is stored and how users can request its deletion. Many data protection laws give users the right to be forgotten, meaning they can request their data be erased. AI agents must adhere to data protection laws like GDPR in the EU or CCPA in California. These laws govern data handling practices and user rights regarding their personal data.
Data Portability and Privacy Policy. Furthermore, developers must create the AI agent’s privacy policy to document and explain to users how their data is handled. This should detail data collection, usage, storage, and user rights. Developers should ensure that they obtain user consent for data collection, especially for sensitive information. Users typically have the option to opt-out or limit the data they provide. In some jurisdictions, users may even have the right to request a copy of their data in a format that can be transferred to another service provider.
Anonymization. For data used in broader analysis or AI training, it should ideally be anonymized to protect individual identities. Developers must understand how their AI agent retrieves and uses historical user data during interactions. This could be for personalization or improving response relevance.
In summary, understanding data privacy for AI agents involves being aware of how user data is collected, used, stored, and protected, and ensuring that users understand their rights regarding accessing, correcting, and deleting their data. Awareness of the mechanisms for data retrieval, both by users and the AI agent, is also crucial for a comprehensive understanding of data privacy.
2.2.4 Interpretability and Explainability
Imitation Learning → Decoupling. Agents are typically trained using a continuous feedback loop in Reinforcement Learning (RL) or Imitation Learning (IL), starting with a randomly initialized policy. However, this approach faces leader-board in obtaining initial rewards in unfamiliar environments, particularly when rewards are sparse or only available at the end of a long-step interaction. Thus, a superior solution is to use an infinite-memory agent trained through IL, which can learn policies from expert data, improving exploration and utilization of unseen environmental space with emergent infrastructure as shown in Fig. 3. With expert characteristics to help the agent explore better and utilize the unseen environmental space. Agent AI, can learn policies and new paradigm flow directly from expert data.
Traditional IL has an agent mimicking an expert demonstrator’s behavior to learn a policy. However, learning the expert policy directly may not always be the best approach, as the agent may not generalize well to unseen situations. To tackle this, we propose learning an agent with in-context prompt or a implicit reward function that captures key aspects of the expert’s behavior, as shown in Fig. 3. This equips the infinite memory agent with physical-world behavior data for task execution, learned from expert demonstrations. It helps overcome existing imitation learning drawbacks like the need for extensive expert data and potential errors in complex tasks. The key idea behind the Agent AI has two parts: 1) the infinite agent that collects physical-world expert demonstrations as state-action pairs and 2) the virtual environment that imitates the agent generator. The imitating agent produces actions that mimic the expert’s behavior, while the agent learns a policy mapping from states to actions by reducing a loss function of the disparity between the expert’s actions and the actions generated by the learned policy.
Decoupling → Generalization. Rather than relying on a task-specific reward function, the agent learns from expert demonstrations, which provide a diverse set of state-action pairs covering various task aspects. The agent then learns a policy that maps states to actions by imitating the expert’s behavior. Decoupling in imitation learning refers to separating the learning process from the task-specific reward function, allowing the policy to generalize across different tasks without explicit reliance on the task-specific reward function. By decoupling, the agent can learn from expert demonstrations and learn a policy that is adaptable to a variety of situations. Decoupling enables transfer learning, where a policy learned in one domain can adapt to others with minimal fine-tuning. By learning a general policy that is not tied to a specific reward function, the agent can leverage the knowledge it acquired in one task to perform well in other related tasks. Since the agent does not rely on a specific reward function, it can adapt to changes in the reward function or environment without the need for significant retraining. This makes the learned policy more robust and generalizable across different environments. Decoupling in this context refers to the separation of two tasks in the learning process: learning the reward function and learning the optimal policy.
Generalization → Emergent Behavior. Generalization explains how emergent properties or behaviors can arise from simpler components or rules. The key idea lies in identifying the basic elements or rules that govern the behavior of the system, such as individual neurons or basic algorithms. Consequently, by observing how these simple components or rules interact with one another. These interactions of these components of ten lead to the emergence of complex behaviors, which are not predictable by examining individual components alone. Generalization across different levels of complexity allows a system to learn general principles applicable across these levels, leading to emergent properties. This enables the system to adapt to new situations, demonstrating the emergence of more com plex behaviors from simpler rules. Furthermore, the ability to generalize across different complexity levels facilitates knowledge transfer from one domain to an other, which contributes to the emergence of complex behaviors in new contexts as the system adapts.
Figure 3: Example of the Emergent Interactive Mechanism using an agent to identify text relevant to the image from candidates. The task involves using a multi-modal AI agent from the web and human-annotated knowledge interaction samples to incorporate external world information.
2.2.5 Inference Augmentation
The inference ability of an AI agent lies in its capacity to interpret, predict, and respond based on its training and input data. While these capabilities are advanced and continually improving, it’s important to recognize their limitations and the influence of the underlying data they are trained on. Particularly, in the context of large language models, it refers to its capacity to draw conclusions, make predictions, and generate responses based on the data it has been trained on and the input it receives. Inference augmentation in AI agents refers to enhancing the AI’s natural inference abilities with additional tools, techniques, or data to improve its performance, accuracy, and utility. This can be particularly important in complex decision-making scenarios or when dealing with nuanced or specialized content. We denote particularly important sources for inference augmentation below:
Data Enrichment. Incorporating additional, often external, data sources to provide more context or background can help the AI agent make more informed inferences, especially in areas where its training data may be limited. For example, AI agents can infer meaning from the context of a conversation or text. They analyze the given information and use it to understand the intent and relevant details of user queries. These models are proficient at recognizing patterns in data. They use this ability to make inferences about language, user behavior, or other relevant phenomena based on the patterns they’ve learned during training.
Algorithm Enhancement. Improving the AI’s underlying algorithms to make better inferences. This could involve using more advanced machine learning models, integrating different types of AI (like combining NLP with image recognition), or updating algorithms to better handle complex tasks. Inference in language models involves understand ing and generating human language. This includes grasping nuances like tone, intent, and the subtleties of different linguistic constructions. Human-in-the-Loop (HITL). Involving human input to augment the AI’s inferences can be particularly useful in areas where human judgment is crucial, such as ethical considerations, creative tasks, or ambiguous scenarios. Humans can provide guidance, correct errors, or offer insights that the agent would not be able to infer on its own. Real-Time Feedback Integration. Using real-time feedback from users or the environment to enhance inferences is another promising method for improving performance during inference. For example, an AI might adjust its recommendations based on live user responses or changing conditions in a dynamic system. Or, if the agent is taking actions in a simulated environment that break certain rules, the agent can be dynamically given feedback to help correct itself. Cross-Domain Knowledge Transfer. Leveraging knowledge or models from one domain to improve inferences in another can be particularly helpful when producing outputs within a specialized discipline. For instance, techniques developed for language translation might be applied to code generation, or insights from medical diagnostics could enhance predictive maintenance in machinery. Customization for Specific Use Cases. Tailoring the AI’s inference capabilities for particular applications or industries can involve training the AI on specialized datasets or fine-tuning its models to better suit specific tasks, such as legal analysis, medical diagnosis, or financial forecasting. Since the particular language or information within one domain can greatly contrast with the language from other domains, it can be beneficial to finetune the agent on domain-specific information. Ethical and Bias Considerations. It is important to ensure that the augmentation process does not introduce new biases or ethical issues. This involves careful consideration of the sources of additional data or the impact of the new inference augmentation algorithms on fairness and transparency. When making inferences, especially about sensitive topics, AI agents must sometimes navigate ethical considerations. This involves avoiding harmful stereotypes, respecting privacy, and ensuring fairness. Continuous Learning and Adaptation. Regularly updating and refining the AI’s capabilities to keep up with new developments, changing data landscapes, and evolving user needs. In summmary, winference augmentation in AI agents involves methods in which their natural inference abilities can be enhanced through additional data, improved algorithms, human input, and other techniques. Depending on the use-case, this augmentation is often essential for dealing with complex tasks and ensuring accuracy in the agent’s outputs. 2.2.6 Regulation Recently, Agent AI has made significant advancements, and its integration into embodied systems has opened new possibilities for interacting with agents via more immersive, dynamic, and engaging experiences. To expedite the process and ease the cumbersome work in agent AI developing, we are proposing to develop the next-generation AI-empowered pipeline for agent interaction. Develop a human-machine collaboration system where humans and machines can communicate and interact meaningfully. The system can leverage the LLM’s or VLM dialog capabilities and vast action to talk with human players and identify human needs. Then it will perform proper actions to help human players upon request. When employing LLM/VLMs for a human-machine collaboration system, it is essential to note that these operate as black boxes, generating unpredictable output. This uncertainty can become crucial in a physical setup, such as operating actual robotics. An approach to address this challenge is constraining the focus of the LLM/VLM through prompt engineering. For instance, in robotic task planning from instructions, providing environmental information within the prompt has been reported to yield more stable outputs than relying solely on text (Gramopadhye and Szafir, 2022). This report is supported by the Minsky’s frame theory of AI (Minsky, 1975), suggesting that the problem space to be solved by LLM/VLMs is defined by the given prompts. Another approach is designing prompts to make LLM/VLMs include explanatory text to allow users understand what the model has focused on or recognized. Additionally, implementing a higher layer that allows for pre-execution verification and modification under human guidance can facilitate the operation of systems working under such guidance (Fig. 4).
Figure 4: A robot teaching system developed in (Wake et al., 2023c). (Left) The system workflow. The process involves three steps: Task planning, where ChatGPT plans robotic tasks from instructions and environmental information; Demonstration, where the user visually demonstrates the action sequence. All the steps are reviewed by the user, and if any step fails or shows deficiencies, the previous steps can be revisited as necessary. (Right) A web application that enables uploading of demonstration data and the interaction between the user and ChatGPT.
2.3 Agent AI for Emergent Abilities
Despite the growing adoption of interactive agent AI systems, the majority of proposed methods still face a challenge in terms of their generalization performance in unseen environments or scenarios. Current modeling practices require developers to prepare large datasets for each domain to finetune/pretrain models; however, this process is costly and even impossible if the domain is new. To address this issue, we build interactive agents that leverage the knowledge-memory of general-purpose foundation models (ChatGPT, Dall-E, GPT-4, etc.) for a novel scenario, specifically for generating a collaboration space between humans and agents. We discover an emergent mechanism— which we name Mixed Reality with Knowledge Inference Interaction—that facilitates collaboration with humans to solve challenging tasks in complex real-world environments and enables the exploration of unseen environments for adaptation to virtual reality. For this mechanism, the agent learns i) micro-reactions in cross-modality: collecting relevant individual knowledge for each interaction task (e.g., understanding unseen scenes) from the explicit web source and by implicitly inferring from the output of pretrained models; ii) macro-behavior in reality-agnostic: improving interactive dimensions and patterns in language and multi-modality domains, and make changes based on characterized roles, certain target variable, influenced diversification of collaborative information in mixed-reality and LLMs. We investigate the task of knowledge-guided interactive synergistic effects to collaborated scene generation with combining various OpenAI models, and show promising results of how the interactive agent system can further boost the large foundation models in our setting. It integrates and improves the depth of generalization, conscious and interpretability of a complex adaptive AI systems.
Figure 5: Our proposed new agent paradigm for a multi-modal generalist agent. There are 5 main modules as shown in the figures: 1) Environment and Perception with task-planning and skill observation; 2) Agent learning; 3) Memory; 4) Agent action; 5) Cognition.
3 Agent AI Paradigm
In this section, we discuss a new paradigm and framework for training Agent AI. We seek to accomplish several goals with our proposed framework:
• Makeuse of existing pre-trained models and pre-training strategies to effectively bootstrap our agents with effective understanding of important modalities, such as text or visual inputs. • Support for sufficient long-term task-planning capabilities. • Incorporate a framework for memory that allows for learned knowledge to be encoded and retrieved later. • Allow for environmental feedback to be used to effectively train the agent to learn which actions to take.
We show a high-level new agent diagram outlining the important submodules of such a system in Fig. 5.
3.1 LLMs and VLMs
We can use the LLM or VLM model to bootstrap the components of the Agent as showed in Fig. 5. In particular, LLMs have been shown to perform well for task-planning (Gong et al., 2023a), contain significant world knowledge (Yu et al., 2023b), and display impressive logical reasoning capabilities (Creswell et al., 2022). Additionally, VLMs such as CLIP (Radford et al., 2021) provide a general visual encoder that is language-aligned, as well as providing zero-shot visual recognition capabilities. For example, state-of-the-art open-source multi-modal models such as LLaVA (Liu et al., 2023c) and Instruct BLIP (Dai et al., 2023) rely upon frozen CLIP models as visual encoders.
3.2 Agent Transformer Definition
Instead of using frozen LLMs and VLMs for the AI agent, it is also possible to use a single-agent transformer model that takes visual tokens and language tokens as input, similar to Gato (Reed et al., 2022). In addition to vision and language, we add a third general type of input, which we denote as agent tokens. Conceptually, agent tokens are used to reserve a specific subspace of the input and output space of the model for agentic behaviors. For robotics or game playing, this may be represented as the input action space of the controller. When training agents to use specific tools, such as image-generation or image-editing models, or for other API calls, agent tokens can also be used. As showed in Fig. 7, we can combine the agent tokens with visual and language tokens to generate a unified interface for training multi-modal agent AI. Compared to using large, proprietary LLMs as agents, there are several advantages to using an agent transformer. Firstly, the model can be easily customized to very specific agentic tasks that may be difficult to represent in natural language (e.g. controller inputs or other specific actions). Thus, the agent can learn from environmental interactions and domain-specific data to improve performance. Secondly, it can be easier to understand why the model does or does not take specific actions by having access to the probabilities of the agent tokens. Thirdly, there are certain domains such as healthcare and law that have strict data privacy requirements. Finally, a relatively smaller agent transformer can potentially be significantly cheaper than a larger proprietary language model.
Figure 6: We show the current paradigm for creating multi-modal AI agents by incorporating a Large Language Model (LLM) with a Large Vision Model (LVM). Generally, these models take visual or language inputs and use pre-trained and frozen visual and language models, learning smaller sub-network that connect and bridge modalities. Examples include Flamingo (Alayrac et al., 2022), BLIP-2 (Li et al., 2023c), InstructBLIP (Dai et al., 2023), and LLaVA (Liu et al., 2023c).Figure 7: The unified agent multi-modal transformer model. Instead of connecting frozen submodules and using existing foundation models as building blocks, we propose a unified and end-to-end training paradigm for agent systems. We can still initialize the submodules with LLMs and LVMs as in Figure 6 but also make use of agent tokens, specialized tokens for training the model to perform agentic behaviors in a specific domain (e.g., robotics). For more details about agent tokens, see Section 3.2
3.3 Agent Transformer Creation
As shown above in Fig. 5, we can use the new agent paradigm with LLM and VLM-bootstrapped agents, as well as leveraging data generated from large foundation models to train the agent transformer model for learning to execute specific goals. Within this process, the agent model is trained to be specialized and tailored for specific tasks and domains. This approach allows you to leverage a pre-existing, foundation model’s learned features and knowledge. We show a simplified overview of the process in two steps below:
Define Objectives within the Domain. In order to train the agent transformer, the objectives and the action-space of the agent within the context of each specific environment needs to be clearly defined. This includes determining which specific tasks or actions the agent needs to perform and assigning unique agent tokens for each. Furthermore, any automatic rules or procedures that can be used to identify successful completion of tasks can significantly improve the amount of data available for training. Otherwise, foundation-model generated or human-annotated data will be required for training the model. After the data is collected and it is possible to evaluate the performance of the agent, the process of continuous improvement can begin.
Continuous Improvement. Continuous monitoring of the model’s performance and collection of feedback are essential steps in the process. Feedback should be used for further fine-tuning and updates. It is also crucial to ensure that the model does not perpetuate biases or unethical outcomes. This necessitates a careful examination of the training data, regular checks for biases in outputs, and, if needed, training the model to recognize and avoid biases. Once the model achieves satisfactory performance, it can be deployed for the intended application. Continuous monitoring remains vital to ensure that the model performs as expected and to facilitate necessary adjustments. More details on this process, sources of training data, and details surrounding continous learning for agent AI can be found in Section 8.
4 Agent AI Learning
4.1 Strategy and Mechanism
The strategy of interactive AI on different domains which extends the paradigm of calling large foundation models with a trained agent that actively seeks to collect user feedback, action information, useful knowledge for generation and interaction. Some times, the LLM/VLM models are not need to trained again, and we improve their performance by providing improved contextual prompts at test time for an agent. On the other hand, it always involves a knowl edge/reasoning/commonsense/inference interactive modeling through a combination of triple systems- one performing knowledge retrieval from multi-model query, second performing interactive generation from the relevant agent, and last one the trained a new, informative self-supervised training or pre-training with reinforcement learning or imitation learning with improved way.
4.1.1 Reinforcement Learning (RL) There is a rich history of leveraging reinforcement learning (RL) to train interactive agents that exhibits intelligent behaviors. RL is a methodology to learn the optimal relationship between states and actions based on rewards (or penalties) received as a result of its actions. RL is a highly scalable framework that has been applied to numerous applications including robotics, however, it generally faces several leader-board and LLM/VLMs have shown their potential to mitigate or overcome some of those difficulties:
• Reward designing The efficiency of policy learning greatly depends on the design of the reward function. Designing the reward function requires not only knowledge of RL algorithms but also a deep understanding of the nature of the task, and thus often necessitates crafting the function based on expert experience. Several studies explored the use of LLM/VLMs for designing reward functions (Yu et al., 2023a; Katara et al., 2023; Maet al., 2023).
• Data collection and efficiency Given its exploratory nature, RL-based policy learning requires a significant amount of data (Padalkar et al., 2023). The necessity for extensive data becomes particularly evident when the policy involves managing long sequences or integrating complex actions. This is because these scenarios demand more nuanced decision-making and learning from a wider range of situations. In recent studies, efforts have been directed towards enhancing data generation to support policy learning (Kumar et al., 2023; Du et al., 2023). Additionally, in some studies, these models have been integrated into the reward function to improve policy learning (Sontakke et al., 2023). Parallel to these developments, another strand of research has focused on achieving parameter efficiency in learning processes using VLMs (Tang et al., 2023; Li et al., 2023d) and LLMs(Shi et al., 2023)
• Long-horizon steps In relation to the issue of data efficiency, RL becomes more challenging as the length of action sequences increases. This is due to the ambiguity in the relationship between actions and rewards, known as the credit assignment problem, and the increase in the number of states to be explored, necessitating a significant amount of time and data. One typical approach for long and complex tasks is to break them down into a sequence of subgoals and apply pretrained policies to solve each subgoal (e.g., (Takamatsu et al., 2022)). This idea falls within the framework called the task and motion planning (TAMP)(Garrett et al., 2021). TAMP is composed of two primary components: task planning, which entails identifying sequences of high-level actions, and motion planning, which involves finding physically consistent, collision-free trajectories to achieve the objectives of the task plan.
LLMsare well-suited to TAMP, and recent research has often adopted an approach where LLMs are used to execute high-level task planning, while low-level controls are addressed with RL-based policies (Xu et al., 2023; Sun et al., 2023a; Li et al., 2023b; Parakh et al., 2023). The advanced capabilities of LLMs enable them to effectively decompose even abstract instructions into subgoals (Wake et al., 2023c), contributing to the enhancement of language understanding abilities in robotic systems.
4.1.2 Imitation Learning (IL)
While RL aims to train a policy based on exploratory behavior and maximizing rewards through interactions with the environment, imitation learning (IL) seeks to leverage expert data to mimic the actions of experienced agents or experts. For example, in robotics, one of the major frameworks based on IL is Behavioral Cloning (BC). BC is an approach where a robot is trained to mimic the actions of an expert by directly copying them. In this approach, the expert’s actions in performing specific tasks are recorded, and the robot is trained to replicate these actions in similar situations. Recent BC-based methods often incorporate technologies from LLM/VLMs, enabling more advanced end-to-end models. For example, Brohan et al. proposed RT-1 (Brohan et al., 2022) and RT-2 (Brohan et al., 2023), transformer-based models that output an action sequence for the base and arm, taking a series of images and language as input. These models are reported to show high generalization performance as the result of training on a large amount of training data.
4.1.3 Traditional RGB
Learning intelligent agent behavior leveraging image inputs has been of interest for many years (Mnih et al., 2015). The inherent challenge of using RGB input is the curse of dimensionality. To solve this problem, researchers either use more data (Jang et al., 2022; Ha et al., 2023) or introduce inductive biases into the model design to improve sample efficiency. In particular, authors incorporate 3D structures into the model architecture for manipulations (Zeng et al., 2021; Shridhar et al., 2023; Goyal et al., 2023; James and Davison, 2022). For robot navigation, authors (Chaplot et al., 2020a,b) leverage maps as a representation. Maps can either be learned from a neural network aggregating all previous RGBinputs or through 3D reconstruction methods such as Neural Radiance Fields (Rosinol et al., 2022). To obtain more data, researchers synthesize synthetic data using graphics simulators (Mu et al., 2021; Gong et al., 2023b), and try to close the sim2real gap (Tobin et al., 2017; Sadeghi and Levine, 2016; Peng et al., 2018). Recently, there has been some collective effort to curate large-scale dataset that aims to resolve the data scarcity problem (Padalkar et al., 2023; Brohan et al., 2023). On the other hand, to improve sample complexity, data augmentation techniques have been extensively studied as well (Zeng et al., 2021; Rao et al., 2020; Haarnoja et al., 2023; Lifshitz et al., 2023).
4.1.4 In-context Learning
In-context learning was shown to be an effective method for solving tasks in NLP with the advent of large language models like GPT-3 (Brown et al., 2020; Min et al., 2022). Few-shot prompts were seen to be an effective way to contextualize model output’s across a variety of tasks in NLP by providing examples of the task within the context of the LLMprompt. Factors like the diversity of examples and quality of examples shown for the in-context demonstrations may improve the quality of model outputs (An et al., 2023; Dong et al., 2022).
Within the context of multi-modal foundation models, models like Flamingo and BLIP-2 (Alayrac et al., 2022; Li et al., 2023c) have been shown to be effective at a variety of visual understanding tasks when given only given a small number of examples. In context learning can be further improved for agents within environments by incorporating environment-specific feedback when certain actions are taken (Gong et al., 2023a).
4.1.5 Optimization in the Agent System
The optimization of agent systems can be divided into spatial and temporal aspects. Spatial optimization considers how agents operate within a physical space to execute tasks. This includes inter-robot coordination, resource allocation, and keeping an organized space.
In order to effectively optimize agent AI systems, especially systems with large numbers of agents acting in parallel, previous works have focused on using large batch reinforcement learning (Shacklett et al., 2023). Since datasets of multi-agent interactions for specific tasks are rare, self-play reinforcement learning enables a team of agents to improve over time. However, this may also lead to very brittle agents that can only work under self-play and not with humans or other independent agents since they over-fit to the self-play training paradigm. To address this issue, we can instead discover a diverse set of conventions (Cui et al., 2023; Sarkar et al., 2023), and train an agent that is aware of a wide range of conventions. Foundation models can further help to establish conventions with humans or other independent agents, enabling smooth coordination with new agents.
Temporal optimization, on the other hand, focuses on how agents execute tasks over time. This encompasses task scheduling, sequencing, and timeline efficiency. For instance, optimizing the trajectory of a robot’s arm is an example of efficiently optimizing movement between consecutive tasks (Zhou et al., 2023c). At the level of task scheduling, methods like LLM-DP (Dagan et al., 2023) and ReAct (Yao et al., 2023a) have been proposed to solve efficient task planning by incorporating environmental factors interactively.
4.2 Agent Systems (zero-shot and few-shot level)
4.2.1 Agent Modules
Our foray into the agent paradigm involves the development of Agent AI “Modules” for interactive multi-modal agents using LLMs or VLMs. Our initial Agent Modules facilitate training or in-context learning and adopt a minimalist design for the purposes of demonstrating the agent’s ability to schedule and coordinate effectively. We also explored initial prompt-based memory techniques that facilitate better planning and inform future actions approaches within the domain. To illustrate, our “MindAgent” infrastructure comprises 5 main modules: 1) environment perception with task planning, 2) agent learning, 3) memory, 4) general agent action prediction and 5) cognition, as shown in Figure 5.
4.2.2 Agent Infrastructure
Agent-based AI is a large and fast-growing community within the domains of entertainment, research, and industry. The development of large foundation models has significantly improved the performance of agent AI systems. However, creating agents in this vein is limited by the increasing effort necessary to create high-quality datasets and overall cost. At Microsoft, building high-quality agent infrastructure has significantly impacted multi-modal agent copilots by using advanced hardware, diverse data sources, and powerful software libraries. As Microsoft continues to push the boundaries of agent technology, AI agent platforms are poised to remain a dominant force in the world of multimodal intelligence for years to come. Nevertheless, agent AI interaction is currently still a complex process that requires a combination of multiple skills. The recent advancements in the space of large generative AI models have the potential to greatly reduce the current high cost and time required for interactive content, both for large studios, as well as empowering smaller independent content creators to design high quality experiences beyond what they are currently capable of. The current human-machine interaction systems inside multi-modal agents are primarily rule-based. They do have intelligent behaviors in response to human/user actions and possess web knowledge to some extent. However, these interactions are often limited by software development costs to enable specific behaviors in the system. In addition, current models are not designed to help human to achieve a goal in the case of users’ inability to achieve specific tasks. Therefore, there is a need for an agent AI system infrastructure to analyze users behaviors and provide proper support when needed.
4.3 Agentic Foundation Models (pretraining and finetune level)
The use of pre-trained foundation models offers a significant advantage in their wide applicability across diverse use cases. The integration of these models enables the development of customized solutions for various applications, circumventing the need for extensive labeled datasets for each specific task. Anotable example in the field of navigation is the LM-Nav system (Shah et al., 2023a), which incorporates GPT-3 and CLIP in a novel approach. It effectively uses textual landmarks generated by the language model, anchoring them in images acquired by robots for navigation. This method demonstrates a seamless fusion of textual and visual data, significantly enhancing the capabilities of robotic navigation, while maintaining wide applicability. In robot manipulation, several studies have proposed the use of off-the-shelf LLMs (e.g., ChatGPT) while using open vocabulary object detectors. The combination of LLM and advanced object detectors (e.g., Detic (Zhou et al., 2022)) fa cilitates the understanding of human instruction while grounding the textual information in scenery information (Parakh et al., 2023). Furthermore, the latest advancements showcase the potential of using prompt engineering with advanced multi-modal models such as GPT-4V(ision) (Wake et al., 2023b). This technique opens avenues for multi-modal task planning, underscoring the versatility and adaptability of pre-trained models in a variety of contexts.
5 Agent AI Categorization
5.1 Generalist Agent Areas
Computer-based action and generalist agents (GAs) are useful for many tasks. Recent progress in the field of large foundation models and interactive AI has enabled new functionalities for GAs. However, for a GA to become truly valuable to its users, it must be natural to interact with, and generalize to a broad range of contexts and modalities. We high-quality extended main Chapters on Agent foundation AI in Sec.6, especially in areas relevant to the themes in general of these topics:
Multimodal Agent AI (MMA) is an upcoming forum(https://multimodalagentai.github.io/) for our research and industry communities to engage with each other and with the broader research and technology communities in Agent AI. Recent progress in the field of large foundation models and interactive AI has enabled new functionalities for generalist agents (GAs), such as predicting user actions and task planning in constrained settings (e.g., MindAgent (Gong et al., 2023a), fine-grained multimodal video understanding (Luo et al., 2022), Robotics (Ahn et al., 2022b; Brohan et al., 2023)), or providing a chat companion for users that incorporates knowledge feedback (e.g., website customer support for healthcare systems (Peng et al., 2023)). More details about the representative works and most recent representative works are shown below. We hope to discuss our vision for the future of MAA and inspire future researchers to work in this space. This article and our forum covers the following main topics, but is not limited exclusively to these:
• Primary Subject Topics: Multimodal Agent AI, General Agent AI • Secondary Subject Topics: Embodied Agents, Action Agents, Language-based Agents, Vision & Language Agents, Knowledge and Inference Agents, Agents for Gaming, Robotics, Healthcare, etc. • Extend Subject Topics: Visual Navigation, Simulation Environments, Rearrangement, Agentic Foundation Models, VR/AR/MR, Embodied Vision & Language.
Next, we present a specific lists of representative agent categories as follows:
5.2 Embodied Agents
Our biological minds live in bodies, and our bodies move through a changing world. The goal of embodied artificial intelligence is to create agents, such as robots, which learn to creatively solve challenging tasks requiring interaction with the environment. While this is a significant challenge, important advances in deep learning and the increasing availability of large datasets like ImageNet have enabled superhuman performance on a variety of AI tasks previously thought intractable. Computer vision, speech recognition and natural language processing have experienced transformative revolutions at passive input-output tasks like language translation and image classification, and reinforcement learning has similarly achieved world-class performance at interactive tasks like game playing. These advances have supercharged embodied AI, enabling a growing collection of users to make rapid progress towards intelligent agents can interactive with machine.
5.2.1 Action Agents
Action agents refer to the agents that need to execute physical actions in the simulated physical environment or real world. In particular, they need to be actively engaging in activities with the environment. We broadly classify action agents into two different categories based on their application domains: gaming AI and robotics. In gaming AI, the agents will interact with the game environment and other independent entities. In these settings, natural language can enable smooth communication between agents and humans. Depending on the game, there may be a specific task to accomplish, providing a true reward signal. For instance, in the competitive Diplomacy game, training a language model using human conversation data along with an action policy with RL enables human-level play (Meta Fundamental AI Research (FAIR) Diplomacy Team et al., 2022).
There are also settings where we agents act as normal residents in a town (Park et al., 2023a), without trying to optimize a specific goal. Foundation models are useful in these settings because they can model interactions that appear more natural by mimicking human behavior. When augmented with external memory, they produce convincing agents that can have conversations, daily schedules, form relationships, and have a virtual life.
5.2.2 Interactive Agents
Interactive agents simply refer to agents that can interact with the world, a broader class of agents than action agents. Their forms of interaction do not necessarily require physical actions, but may involve communicating information to users or modifying the environment. For instance, an embodied interactive agent may answer a user’s questions about a topic through dialogue or help users parse through existing information similar to a chatbot. By extending an agent’s capabilities to include information sharing, the core designs and algorithms of Agent AI can be effectively adapted for a range of applications, such as diagnostic (Lee et al., 2023) and knowledge-retrieval (Peng et al., 2023) agents.
5.3 Simulation and Environments Agents
An effective approach for AI agents to learn how to act in an environment is to go through trial-and-error experiences via interactions with the environment. A representative method is RL, which requires extensive experience of failures to train an agent. Although there exist approaches that use physical agents (Kalashnikov et al., 2018), using physical agents is time-consuming and costly. Furthermore, training in the physical environment is often feasible when failure in actual environments can be dangerous (e.g., autonomous driving, underwater vehicles). Hence, using simulators to learn policies is a common approach.
Many simulation platforms have been proposed for research in embodied AI, ranging from navigation (Tsoi et al., 2022; Deitke et al., 2020; Kolve et al., 2017) to object manipulation (Wang et al., 2023d; Mees et al., 2022; Yang et al., 2023a; Ehsani et al., 2021). One example is Habitat (Savva et al., 2019; Szot et al., 2021), which provides a 3D indoor environment where human- and robotic-agents can perform various tasks such as navigation, instruction following, and question answering. Another representative simulation platform is Virtual Home (Puig et al., 2018), supporting human avatars for object manipulation in 3D indoor environments. In the field of gaming, Carroll et al. have introduced “Overcooked-AI,” a benchmark environment designed to study cooperative tasks between humans and AI (Carroll et al., 2019). Along similar lines, several works aim to incorporate real human intervention beyond the focus of interaction between agents and the environment (Puig et al., 2023; Li et al., 2021a; Srivastava et al., 2022). These simulators contribute to the learning of policies in practical settings involving agent and robot interactions, and IL-based policy learning utilizing human demonstrative actions.
In certain scenarios, the process of learning a policy may necessitate the integration of specialized features within simulators. For example, in the case of learning image-based policies, realistic rendering is often required to facilitate adaptability to real environments (Mittal et al., 2023; Zhong et al., 2023). Utilizing a realistic rendering engine is effective for generating images that reflect various conditions, such as lighting environments. Moreover, simulators employing physics engines are required to simulate physical interactions with objects (Liu and Negrut, 2021). The integration of physics engines in simulation has been shown to facilitate the acquisition of skills that are applicable in real-world scenarios (Saito et al., 2023).
5.4 Generative Agents
The recent advancements in the space of large generative AI models have the potential to greatly reduce the current high cost and time required for interactive content, both for large gaming studios, as well as empower smaller independent studios to create high quality experiences beyond what they are currently capable of. Additionally, embedding large AI models within a sandbox environment will allow users to author their own experiences and express their creativity in ways that are currently out of reach.
The goals of this agent go beyond simply adding interactive 3d content to scenes, but also include:
• Adding arbitrary behavior and rules of interactions to the objects, allowing the user to create their own VR rules with minimal prompting. • Generating whole level geometry from a sketch on a piece of paper, by using the multimodal GPT4-v model, as well as other chains of models involving vision AI models • Retexturing content in scenes using diffusion models • Creating custom shaders and visual special effects from simple user prompts
One potential application in the short term is the VR creation of a storyboarding/prototype tool allowing a single user to create a rough (but functional) sketch of an experience/game an order of magnitude faster than currently feasible. Such a prototype then could be expanded and made more polished using these tools as well.
5.4.1 AR/VR/mixed-reality Agents
AR/VR/mixed-reality (jointly referred to as XR) settings currently require skilled artists and animators to create characters, environments, and objects to be used to model interactions in virtual worlds. This is a costly process that involves concept art, 3D modeling, texturing, rigging, and animation. XR agents can assist in this process by facilitating interactions between creators and building tools to help build the final virtual environment.
Our early experiments have already demonstrated that GPT models can be used in the few-shot regime inside of the Unity engine (without any additional fine-tuning) to call engine-specific methods, use API calls to download 3d models from the internet and place them into the scene, and assign state trees of behavior and animations to them (Huang et al., 2023a). This behavior likely emerges due to the presence of similar code in open source game repositories that use Unity. Therefore, GPT models are capable of building rich visual scenes in terms of loading in many objects into the scene from a simple user prompt.
The aim of this category of agents is to build a platform and a set of tools that provide an efficient interface between large AI models (both GPT-family ones as well as diffusion image models) and a rendering engine. We explore two primary avenues here:
• Integration of large models into the various editor tools in the agent infrastructure, allowing for significant speedups in development. • Controlling the rendering engine from within a user experience, by generating code that follows user instruction and then compiling it at runtime, allowing for users to potentially edit the VR/simulation they are interacting with in arbitrary ways, even by introducing new agent mechanics.
Introducing an AI copilot focused on XR settings would be useful for XR creators, who can use the copilot to complete tedious tasks, like providing simple assets or writing code boilerplate, freeing creators to focus on their creative vision and quickly iterate on ideas.
Furthermore, agents can help users interactively modify the environment by adding new assets, changing the dynamics of the environment, or building new settings. This form of dynamic generation during runtime can also be specified by a creator, enabling the user’s experience to feel fresh and continue evolving over time.
5.5 Knowledge and Logical Inference Agents
The capacity to infer and apply knowledge is a defining feature of human cognition, particularly evident in complex tasks such as logical deduction, and understanding theory of mind(https://plato.stanford.edu/entries/cognitive-science). Making inferences on knowledge ensures that the AI’s responses and actions are consistent with known facts and logical principles. This coherence is a crucial mechanism for maintaining trust and reliability in AI systems, especially in critical applications like medical diagnosis or legal analysis. Here, we introduce agents that incorporate the interplay between knowledge and inference that address specific facets of intelligence and reasoning.
5.5.1 Knowledge Agent Knowledge Agents reason over their acquired knowledge systems in two directions: implicit and explicit. Implicit knowledge is typically what large-scale language models like the GPT series (Brown et al., 2020; OpenAI, 2023) encapsulate after being trained on vast amounts of text data. These models can generate responses that give the impression of understanding, as they draw on patterns and information implicitly learned during training. Explicit knowledge, conversely, is structured and can be directly queried, such as the information found in knowledge bases or databases, which was traditionally used to enhance AI reasoning capabilities by referencing verifiable external resources. Despite the advancements in language models, their implicit knowledge is static and becomes outdated as the world evolves (Lewis et al., 2020; Peng et al., 2023). This limitation necessitates the integration of explicit knowledge sources that are updated continuously, ensuring that AI systems can provide accurate and current responses. The fusion of implicit and explicit knowledge equips AI agents with a more nuanced understanding and the ability to apply knowledge contextually, akin to human intelligence (Gao et al., 2022). Such integration is crucial for crafting knowledge-centric AI agents that not only possess information but can also understand, explain, and employ it, thereby narrowing the chasm between extensive learning and profound knowledge (Marcus and Davis, 2019; Gao et al., 2020). These agents are designed to reason with flexibility and dynamic information about the world, enhancing their robustness and adaptability (Marcus, 2020).
5.5.2 Logic Agents
Generally, a logic agent is a component of a system designed to apply logical reasoning to process data or solve tasks specific to logical inference or logical reasoning. Logic agents within the context of large foundation models like GPT-4 refers to a specialized component or submodules designed to handle logical reasoning tasks. These tasks often involve understanding and manipulating abstract concepts, deducing conclusions from given premises, or solving problems that require a structured, logical approach. Broadly, foundation models like GPT-4 are trained on a vast corpus of text data and learn to perform a wide range of tasks, including those that require some form of logical reasoning. Thus, their capability for logical reasoning is integrated into the overall architecture, and they generally do not possess a distinct, isolated “Logic agent”. While GPT-4 and similar models can perform tasks that involve logic, their approach is fundamentally different from how humans or traditional logic-based systems operate. They do not follow formal logical rules or have an explicit understanding of logic; rather, they generate responses based on patterns learned from the training data. As a result, their performance in logical tasks can be impressive, but it can also be inconsistent or limited by the nature of the training data and the inherent limitations of the model’s design. One example of embedding a separate logical submodule into the architecture is (Wang et al., 2023e), which modifies the token embedding process used by LLMs during pre-training by parsing text into logical segments and explicitly modeling logical hierarchies in the token embeddings.
5.5.3 Agents for Emotional Reasoning
Emotional understanding and empathy are important skills for agents in many human-machine interactions. To illustrate, one important goal for creating engaging dialogue agents is to have the agents act with increased emotion and empathy while minimizing socially inappropriate or offensive outputs. To advance towards this goal for dialogue agents, we released the Neural Image Commenting with Empathy (NICE) dataset (Chen et al., 2021) consisting of almost two million images and the corresponding human-generated comments and a set of human emotion annotations. We also provided a novel pre-training model- Modeling Affect Gneration for Image Comments (MAGIC) (Chen et al., 2021) which aims to generate comments for images, conditioned on linguistic representations that capture style and affect, and to help generate more empathetic, emotional, engaging and socially appropriate comments. Our experiments show that the approach is effective in training a more human-like and engaging image comment agent. Developing empathy-aware agents is a promising direction for interactive agents, and it is important to create agents with emotional understanding capabilities across a wide range of groups and populations, especially considering that many current language models exhibit bias in their emotional understanding and empathetic reasoning capabilities (Mao et al., 2022; Wake et al., 2023d).
5.5.4 Neuro-Symbolic Agents Neuro-Symbolic agents operate on a hybrid system of neurons and symbols (d’Avila Garcez and Lamb, 2020). To solve problems stated in natural language is a challenging task because it requires explicitly capturing discrete symbolic structural information implicit in the input. However, most general neural sequence models do not explicitly capture such structural information, limiting their performance on these tasks. The work (Chen et al., 2020) propose a new encoder-decoder model based on a structured neural representation agent, The encoder of TP-N2F employs TPR ‘binding’ to encode natural-language symbolic structure in vector space and the decoder uses TPR ‘unbinding’ to generate, in symbolic space, a sequential program represented by relational tuples, each consisting of a relation (or operation) and a number of arguments. Instruction following vision-language (VL) models like GPT-4 offer a flexible interface that supports a broad range of multimodal tasks in a zero-shot fashion. However, interfaces that operate on full images do not directly enable the user to “point to” and access specific regions within images. This capability is important not only to support reference-grounded VL benchmarks, but also, for practical applications that require precise within-image reasoning. In (Park et al., 2023b), we build Localized Visual Commonsense model which allows users to specify (multiple) regions-as-input. We train our model by sampling localized commonsense knowledge from a large language model (LLM): specifically, we prompt a LLM to collect common sense knowledge given a global literal image description and a local literal region description automatically generated by a set of VL models. This pipeline is scalable and fully automatic, as no aligned or human-authored image and text pairs are required. With a separately trained critic model that selects high quality examples, we find that training on the localized commonsense corpus expanded solely from images can successfully distill existing VL models to support a reference-as-input interface. Empirical results and human evaluations in zero-shot settings demonstrate that our distillation method results in more precise VL models of reasoning compared to a baseline of passing a generated referring expression.
5.6 LLMsandVLMsAgent A number of works leverage LLMs as agents to perform task planning (Huang et al., 2022a; Wang et al., 2023b; Yao et al., 2023a; Li et al., 2023a), and leverage the LLMs’ large internet-scale domain knowledge and zero-shot planning abilities to perform agentic tasks like planning and reasoning. Recent robotics research also leverages LLMs to perform task planning (Ahn et al., 2022a; Huang et al., 2022b; Liang et al., 2022) by decomposing natural language instruction into a sequence of subtasks, either in the natural language form or in Python code , then using a low-level controller to execute these subtasks. Additionally, (Huang et al., 2022b), (Liang et al., 2022), and (Wang et al., 2023a) also incorporate environmental feedback to improve task performance. There have also been a number of works that demonstrate the ability of general-purpose visually-aligned large language models trained on large-scale text, image, and video data to serve as a foundation for creating multi-modal agents that are embodied and can act in various environments (Baker et al., 2022; Driess et al., 2023; Brohan et al., 2023).
6 Agent AI Application Tasks
6.1 Agents for Gaming
Games provide a unique sandbox to test the agentic behavior of LLMs and VLMs, pushing the boundaries of their collaborative and decision-making abilities. We describe three areas in particular that highlight agent’s abilities to interact with human players and other agents, as well as their ability to take meaningful actions within an environment.
6.1.1 NPC Behavior
In modern gaming systems, the behavior of Non-Player Characters (NPCs) is predominantly dictated by predefined scripts crafted by developers. These scripts encompass a range of reactions and interactions based on various triggers or player actions within the gaming environment. However, this scripted nature often results in predictable or repetitive NPC behavior which fails to evolve in response to player’s actions or the dynamic environment of the game. This rigidity hampers the immersive experience intended in a dynamic gaming environment. Therefore, there is a burgeoning interest in leveraging LLMs to induce autonomy and adaptability in NPC behavior, making interactions more nuanced and engaging. AI-driven NPCs can learn from player behavior, adapt to varying strategies, and provide a more challenging and less predictable gameplay experience. Large Language Models (LLMs) can significantly contribute to evolving NPC behavior in games. By processing vast amounts of text, LLMs can learn patterns and generate responses that are more varied and human-like. They can be utilized to create dynamic dialogue systems, making interactions with NPCs more engaging and less predictable. Furthermore, LLMs can be trained on player feedback and in-game data to continually refine NPC behaviors, making them more attuned to player expectations and game dynamics.
Figure 8: The embodied agent for user interactive gaming action prediction and interactive editing with Minecraft Dungeons gaming sense simulation and generation via GPT-4V.
6.1.2 Human-NPC
Interaction The interaction between human players and NPCs is a crucial aspect of the gaming experience. The conventional interaction paradigm is primarily one-dimensional, with NPCs reacting in a preset manner to player inputs. This limitation stifles the potential for a more organic and enriching interaction, akin to human-human interaction within the virtual realm. The advent of LLM and VLM technologies holds the promise of transforming this paradigm. By employing these technologies, gaming systems can analyze and learn from human behavior to provide more human-like interactions. This not only enhances the realism and engagement of the game but also provides a platform for exploring and understanding human-machine interaction in a controlled yet complex setting.
6.1.3 Agent-based Analysis of Gaming
Gaming is an integral part of daily life, estimated to engage half of the world’s population(https://www.dfcint.com/global-video-game-audience-reaches-3-7-billion/). Additionally, it exhibits a positive impact on mental health(https://news.microsoft.com/source/features/work-life/mind-games-how-gaming-can-play-a-positive-role-in-mental-health/). However, contemporary game systems exhibit a deficiency in interactions with human players since their behaviors are primarily hand-crafted by game developers. These pre-programmed behaviors frequently fail to adapt to players’ needs. Consequently, there exists a need for new AI systems in games that can analyze player behaviors and furnish appropriate support when necessary. Intelligent interactive systems bear the potential to revolutionize how gamers interact with gaming systems in general. NPCs’ interactions with gamers are no longer confined by the restricted rule sets designed by game developers. They have the potential to adapt seamlessly to gamers’ experiences, providing timely feedback to enrich the gaming experience and elevate the synergy of human-machine interaction.
Figure9: GPT-4V can effectively predict the high-level next actions when given the “action history” and a “gaming target” in the prompt. Furthermore, GPT-4V accurately recognized that the player is holding wooden logs in their hand and can incorporate this perceived information into its plan for future actions. Although GPT-4Vappearstobecapable of predicting some low-level actions (such as pressing ‘E‘ to open the inventory), the model’s outputs are not inherently suitable for raw low-level action prediction (including mouse movements) and likely requires supplemental modules for low-level action control.
LLMs can serve as a robust tool for analyzing in-game text data, including chat logs, player feedback, and narrative content. They can help in identifying patterns of player behavior, preferences, and interactions which can be invaluable for game developers to improve game mechanics and narratives. Additionally, VLMs can parse through large quantities of image and video data from gaming sessions to help analyze user intent and actions within the game world. Moreover, LLMs and VLMs can facilitate the development of intelligent agents within games that can communicate with players and other agents in a sophisticated and human-like manner, enhancing the overall gaming experience. Beyond LLMs and VLMs, user input data, provides a promising avenue for creating game-playing agents that model perception, game playing, and game understanding by imitating human players. By incorporating a combination of player interactions and feedback, pixel inputs, and natural language planning and understanding, agent models can assist in the continuous improvement of game dynamics, driving a more player-centric evolution of the gaming environment.
6.1.4 Scene Synthesis for Gaming
Scene synthesis is a vital component in the creation and enhancement of immersive gaming environments. It entails the automatic or semi-automatic generation of three dimensional (3D) scenes and environments within a game. This process includes the generation of terrain, placement of objects, creation of realistic lighting, and sometimes even dynamic weather systems.
Modern games often feature vast, open-world environments. Manually designing these landscapes can be in credibly time-consuming and resource-intensive. Automated terrain generation, often leveraging procedural or AI-driven techniques, can produce complex, realistic landscapes with less manual effort. LLMs and VLMs can utilize the internet scale knowledge to formulate rules to design non-repeating landscapes that are visually impressive and unique. Additionally, LLMs and VLMs can be used to ensure the semantic consistency and variability of generated assets. Placing objects such as buildings, vegetation, and other elements within a scene in a realistic and aesthetically pleasing manner is crucial for immersion.
Figure 10: Masked video prediction on unseen Minecraft videos. From left to right: the original frame, the masked frame, the reconstructed frame, and the reconstructed frame with patches.
VLMs and LLMs can assist in object placement by adhering to predefined or learned rules and aesthetics, thus speeding up the level design process. VLMs and LLMs can be further trained to understand the principles of design and aesthetics, aiding in the procedural generation of content. They can help formulate rules or guidelines that procedural algorithms can follow to generate objects, and scenes that are both visually appealing and contextually appropriate.
Realistic lighting and atmospheric effects are fundamental for creating a believable and engaging gaming environment. Advanced algorithms can simulate natural lighting conditions and dynamic weather effects, enhancing the realism and mood of the scene. LLMs can help develop systems to acheive more realistic lighting and atmospheric effects in several innovative ways. VLMs can analyze vast datasets from real-world lighting and atmospheric conditions to help develop more realistic algorithms for simulating these effects in games. By understanding the patterns and intricacies of natural lighting and weather, these models can contribute to the development of algorithms that mimic reality closely. LLMs and VLMs could also be used to develop systems that adjust lighting and atmospheric effects in real-time based on player actions, game states, or external inputs. They can process natural language commands from players to modify the game environment, providing a more interactive and immersive experience.
6.1.5 Experiments and Results
Zero-shot/Few-shot Learning with LLM or LVM. As we showed in the Fig. 8 and Fig. 9, we used GPT-4V for high-level description and action prediction. Fig. 8 showed some qualitative examples of action description generation and editing with GPT-4V. Agent-enhanced text opens up a novel method of generating 3D scenes with game action priors to help improve the naturalness of the scene. Consequently, GPT-4V generates relevant high-level descriptions that are appropriate for the gaming videos.
Small Agent Pretraining Model. To showcase our agent vision-language architecture, we first study its application in a widely used domain for gaming agents by pretraining on Minecraft data. As shown in Fig. 7, given an input action agent, key frame of video, and corresponding text, a standard encoder-decoder can be employed to convert the agent ac tion and image into action text token and image patch token and then use the agent-vision-language decoder to convert it into a action prediction sentence. The overall architecture is depicted in Fig. 7. We evaluate our approach with several Minecraft demonstrations. The Minecraft video data consists of 5min clips, and we use for pretraining contains 78K videos, and we used 5K videos (6% of pretraining data) for the first round pretraining. We train a 250M parameter model on 16 NVIDIAv100GPUsforonedayandvisualize our model out puts in Fig. 10 and Fig. 11. Fig. 10 shows that our relatively small agent architecture can produce reasonable outputs for Minecraft scenes unseen during training. Fig. 11 showed the model’s predictions compared to the ground truth human player actions indicating potential low-level understanding for our small agent model.
Figure 11: The low-level next step action prediction with the small agent pretraining model in gaming Minecraft scene.
Multi-Agent Infrastructure. As showed in the agent paradigm in Fig. 5, we designed a novel infrastructure for a new gaming scenario called “CuisineWorld” (Gong et al., 2023a). We detail our approach in Fig. 12. Our infrastructure allows for multi-agent collaboration by leveraging GPT-4 as a central planner and works across multiple gaming domains. We investigated our system’s multi-agent planning capabilities, and we deployed the infrastructure into real-world video games to demonstrate its multi-agent and human-AI collaboration effectiveness. Additionally, we presented “Cuisineworld”, a text-based multi-agent collaboration benchmark that provides a new auto-metric Collaboration Score (CoS) to quantify collaboration efficiency. Please refer to the Appendix for more examples and details for gaming description, high-level action prediction, and GPT-4V prompting. We show examples for Bleeding Edge in Fig. 32 and Appendix B, Microsoft Flight Simulator in Fig. 33 and Appendix C, ASSASSIN’s CREED ODYSSEY in Fig. 34 and Appendix D, GEARS of WAR 4 in Fig. 35 and Appendix E, and Starfield in Fig. 36 and Appendix F. We also provide a detailed screenshot of the prompting process for GPT4V used to generate Minecraft examples with Fig. 31 in Appendix A.
6.2 Robotics
Robots are representative agents that necessitate effective interaction with their environment. In this section, we will introduce key elements essential for efficient robotic operation, review research topics where the latest LLM/VLM technologies have been applied, and share findings from our most recent studies.
Visual Motor Control. Visual Motor Control refers to the integration of visual perception and motor action to execute tasks effectively in a robotic system. This integration is paramount as it enables robots to interpret the visual data from their environment and accordingly adjust their motor actions to interact with the environment accurately. For instance, in an assembly line, a robot equipped with visual motor control can perceive the position and orientation of objects and accurately align its manipulator to interact with these objects. This capability is essential for ensuring the precision and effectiveness of robotic operations across a myriad of applications, ranging from industrial automation to assisting the elderly in their daily chores. Moreover, visual motor control facilitates robots in adapting to dynamic environments where the state of the environment may change rapidly, requiring real-time adjustments to motor actions based on visual feedback.
Figure 12: The MindAgent of in-context learning gaming Infrastructure. Planning Skill and Tool Use: The game environment requires diverse planning skills and tool use to complete tasks. It generates relevant game information and converts the game data into a structured text format that the LLMs can process. LLM: The main workhorse of our infrastructure makes decisions, thus serving as a dispatcher for the multi-agent system. Memory History: A storage utility for relevant information. Action Module: Extracts actions from text inputs and converted them into domain-specific language and validates DSLs so that they cause no errors during execution.
Additionally, within the context of safe operation, visual information is crucial for detecting execution errors and confirming the pre- and post-conditions of each robot action. In uncontrolled environments, such as unknown domestic settings, robots are more likely to face unexpected outcomes due to unpredictable factors like changing furniture shapes, varied lighting, and slippage. Executing a pre-planned action plan solely in a feedforward manner can pose significant risks in these settings. Therefore, utilizing visual feedback to continually verify outcomes at each step is key to ensuring robust and reliable operation of robotic systems.
Language Conditioned Manipulation. Language Conditioned Manipulation entails the ability of a robotic system to interpret and execute tasks based on language instructions. This aspect is particularly crucial for creating intuitive and user-friendly interfaces for human-robot interaction. Through natural language commands, users can specify goals and tasks to robots in a manner similar to human-human communication, thereby lowering the barrier to operating robotic systems. In a practical scenario, for instance, a user could instruct a service robot to “pick up the red apple from the table,” and the robot would parse this instruction, identify the referred object and execute the task of picking it up (Wake et al., 2023c). The core challenge lies in developing robust natural language processing and understanding algorithms that can accurately interpret a wide array of instructions, ranging from direct commands to more abstract directives, and enable the robot to convert these instructions into actionable tasks. Furthermore, ensuring that robots can generalize these instructions across diverse tasks and environments is critical for enhancing their versatility and utility in real-world applications. The use of language input to guide robot’s task planning has gained attention in the context of a robot framework called Task and Motion Planning (Garrett et al., 2021).
Skill Optimization. Recent studies highlight the effectiveness of LLMs in robotic task planning. However the optimal execution of tasks, especially those involving physical interactions like grasping, requires a deeper understanding of the environment that goes beyond simply interpreting human instructions. For example, robot grasping necessitates precise contact points (Wake et al., 2023e) and arm posture (Sasabuchi et al., 2021) to efficiently execute subsequent actions. While these elements—precise contact points and arm posture—are intuitive for humans, articulating them through language is challenging. Despite advances in internet-scale VLMs, capturing these nuanced indirect cues from scenes and translating them effectively into robotic skills remains a significant challenge. In response, the robotics community is increasingly focusing on collecting enhanced datasets(e.g., (Wang et al., 2023d; Padalkar et al., 2023)) or developing methodologies for direct skill acquisition from human demonstrations (Wake et al., 2021a). Frameworks including Learning-from-Demonstration and Imitation Learning are leading these developments, playing a crucial role in the optimization of physical skills.
6.2.1 LLM/VLM Agent for Robotics.
Recent research has demonstrated the potential of LLM/VLMs for robotic agents that involve interactions with humans in an environment. Research topics that aim to leverage latest LLM/VLM technologies include:
Multimodal Systems: Recent research has been actively focusing on developing end-to-end systems that incorporate the latest LLM and VLM technologies as encoders for input information. Particularly, there is a significant trend towards modifying these foundation models to process multimodal information. (Jiang et al., 2022; Brohan et al., 2023, 2022; Li et al., 2023d; Ahn et al., 2022b; Shah et al., 2023b; Li et al., 2023e). This adaptation aims to guide robotic actions based on both linguistic instructions and visual cues, thus achieving an effective embodiment.
Task Planning and Skill Training: In contrast to end-to-end systems, Task And Motion Planning (TAMP) based systems first compute a high-level task plan and then achieve them with low-level robot control, known as skills. The advanced language processing abilities of LLMs have demonstrated the capability to interpret instructions and decompose them into robot action steps, greatly advancing task planning technologies (Ni et al., 2023; Li et al., 2023b; Parakh et al., 2023; Wake et al., 2023c). For skill training, several studies have explored the use of LLMs/VLMs for designing reward functions (Yu et al., 2023a; Katara et al., 2023; Ma et al., 2023), generating data to facilitate policy learning (Kumar et al., 2023; Du et al., 2023), or serving as part of a reward function (Sontakke et al., 2023). Together with training frameworks such as RL and IL, these efforts will contribute to the development of efficient robot controllers.
On-site Optimization: Executing long task steps in robotics can be difficult due to unexpected and unpredictable environmental conditions. Therefore, a significant challenge in the field of robotics involves dynamically adapting and refining robotic skills by integrating task plans with real-time environmental data. For instance, (Ahn et al., 2022b) proposed an approach that calculates the feasibility of actions (i.e., affordance) from visual information and compares it with planned tasks. Additionally, there are approaches that focus on enabling LLMs to output the pre-conditions and post-conditions (e.g., states of objects and their interrelationships) of task steps to optimize their execution (Zhou et al., 2023c) and detect pre-condition errors for necessary revisions to the task plan (Raman et al., 2023). These strategies seek to achieve environment-grounded robot execution by integrating environmental information and adjusting the robot’s actions at the task plan or controller level.
Conversation Agents: In creating conversational robots, LLMs can contribute to natural, context-sensitive interactions with humans (Ye et al., 2023a; Wake et al., 2023f). These models process and generate responses that mimic human conversation, allowing robots to participate in meaningful dialogues. Additionally, LLMs play a significant role in the estimation of conceptual (Hensel et al., 2023; Teshima et al., 2022) and emotional attributes (Zhao et al., 2023; Yang et al., 2023b; Wake et al., 2023d) of utterances. Those attributes facilitate the understanding of human intent and meaningful gesture generation, thus contributing to the naturalness and efficacy of human-robot communication.
Navigation Agents: Robot navigation has a long history of research, focusing on core aspects such as map-based path planning and Simultaneous Localization and Mapping (SLAM) for creating environmental maps. These functionalities have become standard in widely used robot middleware like the Robot Operating System (ROS) (Guimarães et al., 2016).
While classic navigation techniques remain prevalent in many robotics applications, they typically rely on static or pre-created maps. Recently, there has been an increased interest in advanced technologies that enable robots to navigate in more challenging environments, leveraging breakthroughs in fields like computer vision and natural language processing. One representative task is object navigation (Chaplot et al., 2020a; Batra et al., 2020; Gervet et al., 2023; Ramakrishnan et al., 2022; Zhang et al., 2021), where robots use object names for navigation instead of map coordinates, requiring the visual grounding of object names in the environment. Furthermore, recent attention has been given to technologies that navigate robots in entirely unfamiliar new environments on a zero-shot basis, on top of foundation models, so-called zero-shot object navigation (Gadre et al., 2023; Dorbala et al., 2023; Cai et al., 2023). Additionally, Vision-Language Navigation (VLN) (Anderson et al., 2018a) is a representative task, where the task involves navigating an agent by natural language instructions in previously unseen, real-world environments (Shah et al., 2023a; Zhou et al., 2023a; Dorbala et al., 2022; Liang et al., 2023; Huang et al., 2023b). VLN interprets sentences rather than object names, such as “go to the bathroom on your left.,” thus it requires a higher functionality to parse input text (Wang et al., 2019). The advent of foundation models contributes to the development of such adaptive, on-the-fly navigation technologies by enhancing the understanding of human language instructions and the visual interpretation of environmental information. More detailed explanations of representative VLN research are provided in 6.2.2.
Figure 13: Overview of the robot teaching system that integrates a ChatGPT-empowered task planner. The process involves two steps: Task planning, where the user employs the task planner to create an action sequence and adjusts the result through feedback as necessary, and Demonstration, where the user visually demonstrates the action sequence to provide information needed for robot operation. The vision system collects visual parameters that will be used for robot execution.
6.2.2 Experiments and Results.
An accumulating body of evidence suggests that recent VLMs and LLMs have promising capabilities for symbolic task planning (e.g., what-to-do). However, each task requires low-level control policy (e.g., how-to-do) to achieve successful interaction between the environment. While reinforcement learning and imitation learning are promising approach to learn policies in a data-driven manner, another promising approach is to obtain the strategy directly from humans through on-site demonstration, an approach called Learning-from-Observation (Wake et al., 2021a; Ikeuchi et al., 0). In this section, we introduce a study where we employ ChatGPT for task planning and enrich the plan by parameterizing it with affordance information to facilitate effective and precise execution (Fig. 13).
The pipeline was composed of two modules: task planning and parameterization. In task planning, the system is fed with language instructions and the description of the working environment. These instructions, along with a predefined set of robot actions and output specifications, are compiled into a comprehensive prompt provided to ChatGPT, which then generates a sequence of decomposed tasks with their textual descriptions (Fig. 13; left pane). Notably, we employ a few-shot approach, meaning ChatGPT is not trained on this task, offering an advantage in applicability as it eliminates the need for hardware-dependent data collection and model training. Additionally, the textual descriptions in the output enable the user to check and adjust the results as necessary, which is a crucial feature for a safe and robust operation. Fig. 14 shows the qualitative results conducted for an agentic simulation on top of VirtualHome (Puig et al., 2018). The results demonstrate a reasonable task plan and its flexibility in adjusting outputs, indicating the broad applicability of our approach.
Figure 14: Example of adjusting an output sequence through auto-generated feedback. We use an open-sourced simulator, VirtualHome for the experiment. Given an instruction “Take the pie on the table and warm it using the stove.,” the task planner plans a sequence of functions that are provided in VirtualHome. If an error in execution is detected, the task planner correct its output based on the auto-generated error message.
While the task planner guarantees coherency between the task sequences, successful operation in reality requires detailed parameters. For example, grasp type is crucial for carrying a container while spilling out the content, such a parameter is often ignored in a simulators (see Fig. 14 in grasping a pie). In our robot system, therefore, users are asked to demonstrate each action visually (Fig. 13; right pane). The tasks had predefined parameters necessary for execution, which our vision system extracts from the videos (Wake et al., 2021b). Notably, our robotic system is not designed for exact replication of human motions (i.e., teleoperation) but rather to handle varying real-world conditions, such as changes in object locations. Hence, the parameters extracted from human demonstrations encompass not precise motion paths but affordance information that dictates effective environmental movement (e.g., waypoints for collision avoidance (Wake et al., 2023a), grasp types (Wake et al., 2023e), and upper-limbs postures (Sasabuchi et al., 2021; Wake et al., 2021a)). The posture of the upper limbs is critical in robots with high degrees of freedom and is designed to assume predictable postures for humans coexisting with the operational robot. The task sequence endowed with affordances is transformed into a sequence of reusable robot skills acquired through reinforcement learning and executed by the robot (Takamatsu et al., 2022).
LLM-empowered task planning can be extended to a more versatile robotic system by integrating it with VLMs. Here, we show an example where we use the GPT-4V(ision) to broaden the aforementioned task planner in a multimodal input context (Fig. 15), a human performs actions that are intended to be replicated by the robot. In this paper, only part of the prompt is shown. The whole prompt is available at microsoft.github.io/GPT4Vision-Robot-Manipulation-Prompts.
This pipeline takes demonstration videos and text, then outputs a sequence of robot actions. A vision analyzer aims to understand the actions performed by humans in the video. We used GPT-4V and provided a prompt to generate text instructions in a style typical of human-to-human communication.Fig. 16 demonstrates how the usage of text input allows user to give feedback on GPT-4V’s recognition results for correction purposes. Such a feature, aiming at improving the accuracy of the recognition results, also enables more robust operation.
Figure 15: Overview of the multimodal task planner that leverages GPT-4V and GPT-4. The system processes video demonstrations and text instructions, generating task plans for robotic execution.Figure 16: Examples of the output of the video analyzer. The five frames are extracted at regular intervals and fed into GPT-4V. We describe the entire pipeline in Section 6.2.2.
Next, the scene analyzer compiles the expected work environment into the text information based on the instructions and the first frame of the video data (or an image of the environment). This environmental information includes a list of object names recognized by GPT-4V, the graspable properties of objects, and the spatial relationships between objects. Although these computational processes are a black box within GPT-4V, the information is output based on the knowledge of GPT-4V and the image/text input. Fig. 17 shows the example outputs of our scene analyzer. As shown in the figure, GPT-4V successfully selects the objects that are related to the manipulation. For example, a table is included in the output when the human is relocating a spam container on the table, while the table is ignored for the fridge opening task. These results suggest that the scene analyzer encodes the scene information with respect to the human’s actions. We prompted GPT-4V to explain the results of the object selection process and the reasons behind those choices. In practice, we found this approach resulted in reasonable outputs. Finally, based on the given text instructions and environmental information, the task planner outputs a sequence of tasks (Wake et al., 2023c).
Figure 17: Examples of the outputs of the scene analyzer that leverages GPT-4V. We describe our entire pipeline in Section 6.2.2.
Embodied Agents for Robotics Navigation. Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. Navigation in 3D environments (Zhu et al., 2017a; Mirowski et al., 2016; Mousavian et al., 2018; Hemachandra et al., 2015) is an essential capability of a mobile intelligent system that functions in the physical world. In the past few years, a plethora of tasks and evaluation protocols (Savva et al., 2017; Kolve et al., 2017; Song et al., 2017; Xia et al., 2018; Anderson et al., 2018a) have been proposed as summarized in (Anderson et al., 2018b). VLN (Anderson et al., 2018a) focuses on language-grounded navigation in the real 3D environment. In order to solve the VLN task, (Anderson et al., 2018a) set up an attention-based sequence-to-sequence baseline model. Then (Wang et al., 2018) introduced a hybrid approach that combines model-free and model-based reinforcement learning (RL) to improve the model’s generalizability. Lastly, (Fried et al., 2018) proposed a speaker-follower model that adopts data augmentation, a panoramic action space and modified beam search for VLN, establishing the current state-of-the-art performance on the Room-to-Room dataset. Extending prior work, we propose a Reinforced Cross-Modal Matching (RCM) for VLN in (Wang et al., 2019). The RCM model is built upon (Fried et al., 2018) but differs in many significant aspects: (1) RCM combines a novel multi-reward RL with imitation learning for VLN while Speaker-Follower models (Fried et al., 2018) only uses supervised learning as in (Anderson et al., 2018a). (2) The RCM reasoning navigator performs cross-modal grounding rather than the temporal attention mechanism on single-modality input. (3) The RCM matching critic is similar to the Speaker in terms of the architecture design, but the former is used to provide the cycle-reconstruction intrinsic reward for both RL and SIL training while the latter is used to augment training data for supervised learning. In (Wang et al., 2019), we study how to address three critical leader-board for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problem. As shown in Fig. 18, we propose a novel Reinforced Cross-Modal Matching approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieved a new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%). Moreover, in (Wang et al., 2019) we introduce a self-supervised imitation learning method for exploration in order to explicitly address the generalization issue, which is a problem not well-studied in prior work. Concurrent to the work, (Thomason et al., 2018; Ke et al., 2019; Ma et al., 2019a,b) studies the VLN tasks from various aspects, and (Nguyen et al., 2018) introduces a variant of the VLN task to f ind objects by requesting language assistance when needed. Note that we are the first to propose to explore unseen environments for the VLN task.
Figure 18: Demonstration of embodied agent for the VLN task (Wang et al., 2019). The instruction, the local visual scene, and the global trajectories in a top-down view is shown. The agent does not have access to the top-down view. Path A is the demonstration path following the instruction. Path B and C are two different paths executed by the agent.
6.3 Healthcare
In healthcare, LLMs and VLMs can act as diagnostic agents, patient care assistants, or even therapy aids, but they come with unique leader-board and responsibilities. With the tremendous potential for AI agents to improve patient care and save lives comes an equally dangerous possibility that their misuse or hasty deployment could endanger thousands or millions of people worldwide. We discuss some of the promising routes for AI agents within the context of healthcare and also discuss some of the key leader-board faced.
Diagnostic Agents. Using LLMs as medical chatbots for patient diagnosis has recently attracted great attention due to the high-demand for medical experts and the potential for LLMs to help triage and diagnose patients (Lee et al., 2023). Dialogue agents, especially those that can effectively communicate important medical information to a broad range of people from diverse patient populations, have the potential to provide equitable healthcare access to historically disadvantaged or marginalized groups. Furthermore, doctors and healthcare systems across the world are largely over-burdened and under-resourced, resulting in insufficient access to medical care for hundreds of millions of people worldwide (World Health Organization and World Bank, 2015). Diagnostic agents provide a particularly advantageous pathway to improve healthcare for millions since they have they can be built with the capability to understand a variety of languages, cultures, and health conditions. Initial results have shown that healthcare-knowledgeable LMMs can be trained by utilizing large-scale web data (Li et al., 2023f). Although an exciting direction, the promise of diagnostic agents does not come without risks. We highlight the risks of hallucination within medical contexts, as well as potential pathways for solutions in the following section.
Knowledge Retrieval Agents. Within the medical context, model hallucinations are particularly dangerous and may even result in serious patient harm or death, depending on the severity of the error. For instance, if a patient mistakenly receives a diagnosis suggesting they are free of a condition they actually have, it can lead to catastrophic outcomes. These include postponed or inappropriate treatments, or in some cases, a total lack of necessary medical intervention. The gravity of undiagnosed or misdiagnosed conditions can lead to escalated healthcare expenses, extended therapies causing further physical strain, and in extreme scenarios, severe harm or even death. Thus, approaches that can use agents to more reliably retrieve knowledge (Peng et al., 2023) or generate text in a retrieval-based manner (Guu et al., 2020) are promising directions. Pairing a diagnostic agent with a medical knowledge retrieval agent has the potential to significantly reduce hallucinations while simultaneously improving the quality and preciseness of the responses of the diagnostic dialogue agent.
Telemedicine and Remote Monitoring. Agent-based AI also has great potential within the world of Telemedicine and Remote Monitoring by improving the access to healthcare, improving communications between healthcare providers and patients, as well as improving the efficiency and reducing the costs of frequent doctor-patient interactions (Amjad et al., 2023). Primary care clinicians spend significant amounts of time sifting through patient messages, reports, and emails that are often irrelevant or unnecessary for them to view. There is significant potential to allow for support agents to help triage messages from doctors, patients, and other healthcare providers and to help highlight important messages for all parties. By enabling agentic AI systems to coordinate with patients, clinicians, and other AI agents, there is a massive potential to revolutionize the remote healthcare and digital health industry.
6.3.1 Current Healthcare Capabilities
Image understanding. We demonstrate the current capabilities and limitations of modern multimodal agents such as GPT-4V within the context of healthcare in Fig. 19. We can see that although GPT-4V possesses significant internal knowledge of the equipment and procedures involved in hospital care, it does not always respond to more prescriptive or diagnostic queries by the user.
Video understanding. We investigate the performance of VLM agents for medical video understanding in two contexts. First, we investigate the ability for VLM agents to identify important patient care activities in clinical spaces. Secondly, we explore the usage of of VLMs for more technical videos such as ultrasounds. Specifically, in Figure 20, we demonstrate some of the current capabilities and limitations of GPT-4V for hospital care and medical video analysis.
6.4 Multimodal Agents
The integration of visual and linguistic understanding is crucial for developing sophisticated multimodal AI agents. This includes tasks such as image captioning, visual question answering, video language generation, and video understanding, amongst others. We aim to delve into these visual-language tasks, exploring the leader-board and opportunities they present in the context of AI agents.
6.4.1 Image-Language Understanding and Generation
Image-language understanding is a task that involves the interpretation of visual content in a given image with language and the generation of associated linguistic descriptions. This task is critical to the development of AI agents that can interact with the world in a more human-like manner. Some of most popular ones are image captioning (Lin et al., 2014; Sharma et al., 2018; Young et al., 2014; Krishna et al., 2016), referring expression (Yu et al., 2016; Karpathy et al., 2014), and visual question answering (Antol et al., 2015; Ren et al., 2015; Singh et al., 2019).
More recently, knowledge-intensive Visual Question Answering tasks such as OKVQA (Marino et al., 2019), KB VQA(Wangetal.,2015), FVQA(Wangetal.,2017), and Web QA(Changetal.,2021)have been introduced. Multimodal agents should capable of identifying objects in an image, comprehending their spatial relationships, generating accurate descriptive sentences about the scene, and utilizing reasoning skills to handle knowledge-intensive visual reasoning. This requires not just object recognition capabilities, but also a deep understanding of spatial relationships, visual semantics, and the ability to map these visual elements to linguistic constructs with integration of the world knowledge.
Figure 19: Example prompts and responses when using GPT-4V within the domain of healthcare image understanding. From left to right: (1) an image of a nurse and doctor conducting a CT scan, (2) a synthetic image of an irregular EKG scan, and (3) an image from the ISIC (Codella et al., 2018) skin lesion dataset. We can see that GPT-4V possesses significant medical knowledge and is able to reason about medical images. However, due to safety training, it is unable to make diagnoses for some medical images.
6.4.2 Video and Language Understanding and Generation
Video-language generation. Video captioning or video storytelling is the task of generating a sequence of coherent sentences for a stream of video frames. Inspired by the successful use of recurrent large foundation models employed in video and language tasks, variants of agent driven enhanced models have shown promising results on the task of video-lanaguage generation. The fundamental challenge is that the strong performance of neural encoder-decoder models does not generalize well for visual storytelling, because the task requires a full understanding of the content of each image as well as the relation among different frames. One important goal for the field is to create an agent-aware text-synthesis model that can efficiently encode the sequence of frames and generate a topically coherent multi-sentence paragraph.
Video Understanding. Video understanding extends the scope of image understanding to dynamic visual content. This involves interpretation and reasoning about the sequence of frames in a video, often in conjunction with accompanying audio or textual information. An agent should be able interact with various modalities from visual, text, and also audio modalities to demonstrate their advanced comprehension of video content. Tasks in this domain include video captioning, video question answering, and activity recognition, amongst others. The leader-board in video understanding are manifold. They include the temporal alignment of visual and linguistic content, the handling of long sequences of frames, and the interpretation of complex activities that unfold over time. Regarding audio, the agent could process spoken words, background noises, music, and tone of voice to comprehend the mood, setting, and subtleties of the video content.
Figure 20: Example prompts and responses when using GPT-4V within the domain of healthcare video understanding. Weinput the example videos as 2×2 grids with overlaid text indicating the order of frames. In the first two examples, we prompt GPT-4V to examine the frames in the video to detect the clinical bedside activities performed on the volunteer patients. For the final example, we attempt to prompt GPT-4V to assess an echo cardiogram video, however due to GPT-4V’s safety training, it does not provide a detailed response. For clarity, we bold text that describes the activity of interest, and abbreviate model responses that are unnecessary. We gray-out faces from the individuals to preserve their privacy.Figure 21: Interactive multimodal agents include four main pillars: Interaction, Speech, Vision, and Language. Co-pilot agents are made up of different services. 1) Interaction services help make a unified platform for automated actions, cognition, and decision-making. 2) Audio services integrate audio and speech processing into apps and services. 3) Vision services identify and analyze content within images, videos, and digital ink. 4) Language services extract meaning from structured and unstructured text.
Previous works have focused on employing existing video-language training data available online for establishing video foundational models (Li et al., 2020, 2021b; Fu et al., 2022; Bain et al., 2021; Zellers et al., 2021, 2022; Fu et al., 2023). Supporting such training pipelines and functionalities is, however, difficult due to the limited and often inconsistent nature of these datasets. Video foundational models are designed with masked and contrastive pretraining objectives and later tuned on their respective tasks. Despite showing remarkable results in multimodal benchmarks, these models encounter difficulties in video-only tasks such as action recognition due to their dependency on limited video-text data built from noisy audio transcriptions. This limitation also leads to the lack of robustness and fine-grained reasoning skills that large language models generally possess.
Other methods, similar to those used in image-language understanding, have drawn on the strong reasoning skills and broad knowledge of large language models to improve different facets of video interpretation. The task of video understanding is simplified by language only models like ChatGPT and GPT4 or image-language models like GPT4-V, which treat the audio, video, and language modalities as individual interpretable input data types and position the agents as strong open-source models. For example, (Huang et al., 2023c; Li et al., 2023g) transformed video understanding into a natural language processing (NLP) question-answering formulation by textualizing video content with open-source vision classification/detection/caption models. (Lin et al., 2023) integrated GPT4-V with specialized tools in vision, audio, and speech, to facilitate complex video understanding tasks, such as scripting character movements and actions in long-form videos.
Parallel research explores generating scaled datasets from large models, then applying visual instruction tuning (Liu et al., 2023c; Li et al., 2023c; Zhu et al., 2023) on the generated data. Considerable audio, speech, and visual expert perception models are subsequently used to verbalize videos. Speech is transcribed with automatic speech recognition tools, and video descriptions and related data are produced with various tagging, grounding, and captioning models (Li et al., 2023g; Maaz et al., 2023; Chen et al., 2023; Wang et al., 2023f). These techniques demonstrate how instruction tuning video-language models on generated datasets may lead to enhanced video-reasoning and communication abilities.
6.4.3 Experiments and Results
• Knowledge-Intensive Models: As introduced in INK (Park et al., 2022), and KAT (Gui et al., 2022a), an intensive neural knowledge task that incorporates required knowledge annotated by humans to support knowledge-intensive retrieval task. • Multimodal-Agents: There has been a growing interest in multimodal language models like Chameleon (Lu et al., 2023) and MM-React (Yang et al., 2023c). • Visual Instruction Tuning: VCL(Gui et al., 2022b), Mini-GPT4 (Zhu et al., 2023), MPLUG-OWL (Ye et al., 2023b), LSKD (Park et al., 2023c) generate image-level instruction tuning dataset.
Knowledge-Intensive Agent. As showed in Fig. 22 and Fig. 23, Knowledge-based visual question answering and vision-language retrieval tasks are challenging tasks in multi-modal machine learning that requires outside knowledge beyond image contents. Recent studies on large-scale transformers have primarily focused on maximizing the efficiency of the model’s parameters to store information. This line of research explores a different aspect: whether multimodal transformers can use explicit knowledge in their decision-making process. Pretraining methods based on transformers have shown remarkable success in implicitly learning knowledge representations across multiple modalities. However, traditional methods, mainly unimodal, have investigated knowledge retrieval and subsequent answer prediction, raising questions about the quality and relevance of the knowledge retrieved and the integration of reasoning processes using both implicit and explicit knowledge. To tackle these issues, we introduce the Knowledge Augmented Transformer (KAT), which outperforms others by 6% on the 2022 OK-VQA open-domain multimodal task. KAT combines implicit knowledge from GPT3 with explicit knowledge from websites using an encoder-decoder structure, and allows for concurrent reasoning with both knowledge types during answer generation. Furthermore, incorporating explicit knowledge enhances the interpretability of the model’s predictions. The code and pre-trained models are available at https://github.com/guilk/KAT.
Vision-language Transformer Agent. Next, we introduce the “Training Vision-Language Transformers from Cap tions” (VLC) model (Gui et al., 2022b), a transformer that has been pretrained exclusively with image-caption pairs. Despite using just a simple linear projection layer for image embeddings, VLC attains competitive results across various vision-language tasks, in contrast to other methods that depend on object detectors or supervised CNN/ViT networks.
Figure 22: Example of Intensive Neural Knowledge (INK) (Park et al., 2022) task that uses knowledge to identify text relevant to the image from a set of text candidates. Our task involves leveraging visual and text knowledge retrieved from web and human-annotated knowledge.Figure 23: The KAT model (Gui et al., 2022a) uses a contrastive-learning-based module to retrieve knowledge entries from an explicit knowledge base and uses GPT-3 to retrieve implicit knowledge with supporting evidence. The integration of knowledge is processed by the respective encoder transformer and jointly with reasoning module and the decoder transformer via end-to-end training for answer generation.Figure 24: The overall architecture of the VLC model (Gui et al., 2022b). Our model consists of three modules: (1) Modality-specific projection. We use a simple linear projection to embed patched images and a word embedding layer to embed tokenized text; (2) Multi-modal encoder. We use a 12-layer ViT (Dosovitskiy et al., 2021) initialized from MAE(Heet al., 2022) (ImageNet-1K without labels) as our backbone; (3) Task-specific decoder. We learn our multi-modal representations by masked image/language modeling and image-text matching which are only used during pre-training. We use a 2-layer MLP to fine-tune our multi-modal encoder for downstream tasks. Importantly, we find that the masked image modeling objective is important throughout second-stage pre-training, not only for initialization of the visual transformer.
Through extensive analysis, we explore the potential of VLC as a vision-language transformer agent. For instance, we show that VLC’s visual representations are highly effective for ImageNet-1K classification, and our visualizations confirm that VLC can accurately match image patches to corresponding text tokens. The scalability of performance with more training data highlights the promising potential for developing large-scale, weakly-supervised, open-domain vision-language models.
6.5 Video-language Experiments
To understand the practicality of converting pre-trained image-LLMs for video understanding, we temporally expand and fine-tune Instruct BLIP (Dai et al., 2023) for video captioning. Specifically, we expand the visual encoder of Instruct BLIP (EVA-CLIP-G (Sun et al., 2023b)) using the same divided space-time attention scheme as Frozen in Time (Bain et al., 2021) and keep the Q-former and LLM (Flan-T5-XL (Chung et al., 2022)) frozen during training. We freeze all spatial layers of the visual encoder, while keeping the temporal layers unfrozen during captioning training. This allows for our model to take image and videos as input (matching the image-level performance of Instruct BLIP). We train on a 5 million video-caption subset of WebVid10M (Bain et al., 2021). We visualize two example outputs in Figure 25. However, existing agents fail to fully comprehend precise, fine-grained visual details in the video content. A similar limitation is seen by visual instruction tuning methods, where they lack the general, human-level perception abilities that are remain to be solved by multimodal models and agents.
The instruction-tuned models show promise in accurately summarizing visible actions within videos and identifying actions like “person sitting on a bench” effectively in Fig. 25. However, they sometimes add incorrect details, such as “person smiling to the camera,” revealing a shortfall in capturing conversation topics or the video’s ambiance, elements that are readily apparent to human observers. This shortfall underscores another key limitation: the omission of audio and speech modalities that would enrich the video understanding with context, aiding in more accurate interpretation and preventing such misrepresentations. Bridging this gap requires a holistic integration of available modalities, allowing multimodal agents to reach a level of comprehension akin to human perception and ensuring a fully multimodal approach to video interpretation.
Figure 25: Example prompts and responses when using a video fine-tuned variant of InstructBLIP (method described in Section 6.5). Our model is able to produce long-form textual responses that describe scenes and is able to answer questions related to the temporality of events in the videos.Figure 26: The audio-multimodal agent described in Section 6.5. Hallucinated content are highlighted in red. We use GPT-4V to generate 1) the videochat summary with video frames; 2) the video summary with the frame captions; 3) the video summary with frame captioning and audio information.Figure 27: An interactive multimodal agent that incorporates visual, audio, and text modalities for video understanding. Our pipeline mines hard negative hallucinations to produce difficult queries for the VideoAnalytica challenge. More the related details of interactive audio-video-language agent dataset are described in Section 9.2.
Audio-Video-Language Agents with GPT-4V. We then evaluate the capabilities of GPT-4V as a multimodal agent that integrates vision, audio, and speech for a nuanced and precise understanding of videos, following the methodology outlined in (Lin et al., 2023). Results depicted in Fig. 26 compare the performance of various video agents on the task of video summarization. The video-instruction tuned model (Li et al., 2023g) provides accurate content but falls short on comprehensiveness and detail, missing specific actions like the methodical use of a broomstick to measure a tree’s height.
To enhance the accuracy of video descriptions, we employ GPT-4V to caption frames, while audio and its transcriptions are sourced from the OpenAI Whisper model. We then prompt GPT-4V to create video summaries using only frame captions and then using both frame captions and audio transcriptions. Initially, we observe that frame captions alone can lead to fabricated events, such as a person biting down on a stick in the third segment. These inaccuracies persist in the video summary, with descriptions like “in a playful twist, he bites down on it while holding it horizontally.” Without audio input, the agent cannot correct these captioning errors, resulting in descriptions that are semantically correct but visually misleading.
However, when we provide the audio transcriptions to the agent, it manages to accurately depict the content, even capturing detailed physical actions like “holding the broomstick perpendicular to the body and rotating it downwards.” This level of detail is significantly more informative and gives viewers a clearer understanding of the video’s purpose and key details. These findings highlight the importance of integrating audio, video, and language interactions to develop high-quality multimodal agents. GPT-4V emerges as a promising foundation for such advanced multimodal understanding and interaction.
Embodied Multi-modal Agents with GPT-4V. As shown in Fig. 27, We mainly used StackOverflow to get the initial Question, then we used the “Bing search” API to retrieve a related video and audio corresponding to the question. Next, we mainly use GPT-4V to get the relevant text information and high-level video description. On the other hand, we transfer the key frame audio to a low-level segment description of the key frames via ASR. Finally, we use GPT-4V to generate convincing “hallucinations” that serve as hard negative queries for video-question and answer tasks. We support interactions and question answering in the current frame of the video, as well as summarization for the overall high-level video description. During inference, we also combine external knowledge information via web search to improve answering capapbilities.
The main prompt information for GPT-4V is described as below. The entire prompt is indented for clarity; it is over one page long.
GPT-4V are an assistant to provide descriptive, informative, and full comprehensive details in the video for the visually impaired who can hear the video but cannot see. The job is to create high-quality, dense descriptions of the video by synthesizing the given annotations and output them as JSON. Specifically, GPT-4V will be given original query used to search the video, the video title, description, audio transcription, and potentially noisy descriptions for specific time in the video. Different segments of same video is annotated as “[time start- time end (in seconds)] ’text’ “. Utilize the transcriptions and descriptions all together to reason about the exact detail and visual demonstration that might be happening in the video. GPT-4V will to combine or segment the timestamps as necessary to provide the best segmentation of the video.
Expectations for GPT-4V Output:
1. Action-Oriented Descriptions: Prioritize plausible actions, motions, and physical demonstrations that the audio implies, enriching your narrative with dynamic visual cues.
2. Complete Video Coverage: Provide a continuous and consistent audio-descriptive experience that covers every moment of the video’s duration, ensuring no content is left undescribed.
3. Concise Segmentation: Construct your descriptions in focused, succinct segments of 1-2 sentences each to effectively communicate visual actions without overwhelming detail.
4. Contextual Audio-Visual Synthesis: Seamlessly blend the spoken audio content with inferred visual elements to form a narrative that reflects potential onscreen activities.
5. Imaginative and Plausible Speculation: Infuse your descriptions with creative yet believable visual details that correspond with the audio, enhancing scene comprehension.
6. Accurate Timecode Correspondence: Align your descriptive segments with corresponding time codes, ensuring that speculative visual details synchronize with the audio narrative’s timeline.
7. Confident Narrative Delivery: Present the descriptions with assurance, as though the speculated visuals are occurring, to instill confidence in the listener.
8. Omit Implausible Details: Exclude descriptions of objects or events that do not reasonably fit within the context established by the audio and visual information provided.
The final output should be structured in a JSON format containing a list of dictionaries, each detailing a segment of the video.
The final output should be structured in a JSON format containing a list of dictionaries, each detailing a segment of the video.
For MCCreation: our task is to create multiple-choice questions for video-to-text retrieval tasks that is trivially solved by looking at the title and reading through audio transcriptions. To do so, we will be given original query to get the video, description, audio transcription, and potentially noisy descriptions for specific time in the video.
• Format of audio transcription:-[start-end time in seconds] “transcription” • Format of noisy description:- [time in seconds] “description”
We kindly ask GPT-4V to generate four queries, where the primary query is aligned with the video content, and the other three negatives are subtly different from our primary one. Selecting the primary one should not simply involve listening to audio transcriptions e.g. the text original query is contained in audio transcriptions. The negatives should be closely related but not fully aligned with the video content, requiring visual understanding of the video to differentiate. For example, modify the semantics in nuanced way so that one needs to watch the video than just listening to select the original query. Compile four queries in caption-like statement, with the first one being the rephrased original.
Think step by step how you can come up with negative statements using the information from the video. And justify the negative queries are incorrect but still compelling choices that demand nuanced understanding of the video. And how humans would not accidentally choose the negatives over the original query. Finally, we present the work in the following format of analyses and 4 queries. No need to generate how you translated the original query.
Recognizing task directives and taking action has been a fundamental challenge in interactive AI and natural language processing for decades. With the recent advances in deep learning, there is a growing interest in studying these areas jointly to improve human-agent collaboration. We identify three specific directions, among others, to improve language-grounded agents:
• Tool use and querying from knowledge bases. This direction emphasizes the importance of integrating external knowledge bases, web search, or other helpful tools into the reasoning processes of AI agents. By leveraging structured and unstructured data from various sources, agents can enhance their understanding and provide more accurate and context-aware responses. Furthermore, it fosters the agent’s ability to proactively seek out information when faced with unfamiliar scenarios or queries, ensuring more comprehensive and informed responses. Examples include Toolformer (Schick et al., 2023) and Retrieve What You Need (Wang et al., 2023g).
• Improved agent reasoning and planning. Enhancing the agent’s ability to reason and plan is pivotal for effective human-agent collaboration. This involves the development of models that can understand complex instructions, infer user intentions, and predict potential future scenarios. This can be accomplished by asking the agent to reflect on past actions and failures as in ReAct (Yao et al., 2023a), or by structuring the agent thought process as a form of search (Yao et al., 2023b). By simulating different outcomes and assessing the ramifications of various actions, agents can make more informed context-aware decisions.
• Incorporating system and human feedback. AI agents can frequently operate in two primary contexts: environments that provide explicit signals about the effectiveness of their actions (system feedback), and settings where they collaborate with humans who can offer verbal critiques (human feedback). This direction underscores the need for adaptive learning mechanisms that allow agents to refine their strategies and rectify mistakes, such as in AutoGen (Wu et al., 2023). The ability to continuously learn and adapt from diverse feedback sources ensures that agents remain helpful and aligned for user needs.
6.6.2 General LLM agent
Recognizing and understanding agent content and natural language has been a fundamental challenge in interactive AI and natural language processing for decades. With the recent advances in deep learning, there is a growing interest in studying these two areas jointly for deep understanding of both agent planning or human feedback for knowledge inference and natural language generation. These are the key components of many human-machine-interaction agents, such as “AutoGen”(Wu et al., 2023) and “Retrieve What You Need”(Wang et al., 2023g).
Figure 28: The training recipe used to train the Alpaca model (Taori et al., 2023). At a high level, existing LLMs are used to generate a large pool of instruction-following examples from a smaller set of seed tasks. The generated instruction-following examples are then used to instruction-tune an LLM where the underlying model weights are available.
6.6.3 Instruction-following LLM agents
Furthermore, the creation of LLM Agents that can be trained to effectively follow human instructions has become an important area of research. Initial models used human feedback to train a proxy reward model to simulate human preferences, through a process known as Reinforcement Learning with Human Feedback (RLHF) (Ouyang et al., 2022). This process produced models such as InstructGPT and ChatGPT. In order to more efficiently train instruction-following LLMagents without needing human labels, researchers developed a more efficient method for instruction-tuning that trains the LLM agent directly on instruction/response pairs, either generated by humans like Dolly 2.0(Dolly 2.0 blogpost link) or automatically from LLMs like Alpaca (Taori et al., 2023). We show the overall Alpaca training pipeline in Figure 28.
6.6.4 Experiments and Results
Despite the growing adoption of conversational and self-feedback systems, these forms of AI still do not perform well with regard to generating factually correct responses from their own implicit knowledge and therefore often use external tools like web search and knowledge retrieval mechanisms at inference-time to augment their response as a consequence. Addressing this would help create more engaging experiences for users in many real-life applications. In social conversations (such as those on social media platforms like Instagram and Facebook), or with Q+A websites (such as Ask or Quora), people usually engage with others through a series of comments and by web-searching for information and knowledge relevant to the discussion. Thus, the task of generating conversational turns in this context is not to simply bootstrap upon traditional NLP models and tasks, but to use agents to generate dialogue through intelligent behaviors that reflect knowledge search and acquisition (Peng et al., 2023). In this way, intelligent agents for NLP tasks extends the task description and improves upon the interpretability of the response by adding an explicit knowledge search and retrieval step during dialogue. Incorporating these web search and retrieval agents as feedback during dialogue will help to engage further and deeper the social interactions between humans and agents (Wang et al., 2023e). As the Fig 29 showed, we introduced a new modeling paradigm for transformer language models that detects and extracts important logical structures and information from input texts and then integrates them into the input embeddings through carefully designed multi-layer hierarchical logical projections to infuse logical structures into pre-trained language models as one kind of NLP agent. (Wang et al., 2023e) propose a novel approach to construct logic-aware input embeddings for transformer language models through a combination of logic detection, logic mapping and hierarchical logical projections, and then develop a corresponding new modeling paradigm that can upgrade all existing transformer language models into logical transformers to consistently boost their performance. The proposed logical transformer agent consistently achieve superior performance over their baseline transformer models through a deeper understanding of the logical structures of texts. To human users, it is often these aspects that are more important for delivering a meaningful and interesting conversation via a agent-based coordination between dialogue and information retrieval. Delving deep into natural language processing, this topic will discuss the advancements and leader-board in making LLMs more agentic and better suited for various language-centered tasks.
Figure 29: The logic transformer agent model (Wang et al., 2023e). We integrate a logical reasoning module into the transformer-based abstractive summarization model in order to endow the logic agent the ability to reason over text and dialogue logic, so that it can generate better-quality abstractive summarizations and reduce factuality errors.
An open-domain question answering (QA) system usually follows a retrieve-then-read paradigm, in which a retriever is used to retrieve relevant passages from a large corpus, and then a reader generates answers based on the retrieved passages and the original question. In (Wang et al., 2023g), we propose a simple and novel mutual learning framework to improve the performance of retrieve-then-read-style models via an intermediate module named the knowledge selector agent, which we train with reinforcement learning. The fine-grained knowledge selector into the retrieve-then reader paradigm, whose goal is to construct a small subset of passages which retain question-relevant information. As showed in Figure 30, The knowledge selector agent is trained as a component of our novel mutual learning framework, which iteratively trains the knowledge selector and the reader. We adopt a simple and novel approach employing policy gradients to optimize the knowledge selector agnet, using feedback from the reader to train it to select a small and informative set of passages. This approach avoids brute-force search or manually-designed heuristics, without requiring any annotated query-document pairs for supervision. We show that iteratively training the reader and the knowledge selector agent leads to better predictive performance on some public open-domain question answering benchmarks.
Figure 30: Architecture of one proposed NLP agent (Wang et al., 2023g) mutual learning framework. In each epoch, Phase 1 and Phase 2 are executed alternately. During Phase 1, the parameters of the reader model remain fixed, and only the weights of the knowledge selector are updated. Conversely, during Phase 2, the reader model’s parameters are adjusted, while the knowledge selector’s weights remain frozen.
7 AgentAIAcross Modalities, Domains, and Realities
7.1 Agents for Cross-modal Understanding
Multi-modal understanding is a significant challenge for creating generalist AI agents due to the lack of large-scale datasets that contain vision, language, and agent behavior. More generally, training data for AI agents is often modality specific. This results in most modern multi-modal systems using a combination of frozen submodules. Some notable examples are Flamingo (Alayrac et al., 2022), BLIP-2 (Li et al., 2023c), and LLaVA (Liu et al., 2023c), all of which utilize a frozen LLM and frozen visual encoder. These submodules are trained individually on separate datasets, and then adaptation layers are trained to encode the visual encoder into the LLM embedding space. In order to make further progress for cross-modal understanding for AI agents, it is likely that the strategy of using frozen LLMs and visual encoders will need to change. Indeed, RT-2, a recent visual-language model that is capable of taking actions within the domain of robotics showed significantly improved performance when jointly tuning the visual encoder and LLM for robotics and visual-language tasks (Brohan et al., 2023). 7.2 Agents for Cross-domain Understanding Akey challenge for creating generalist agents is the distinctive visual appearance and disparate action spaces across different domains. Humans possess the capability to interpret images and videos from various sources, including the real world, video games, and specialized domains such as robotics and healthcare, once they become familiar with the specific details of these areas. However, existing LLMs and VLMs often demonstrate significant differences between the data they were trained on and the varied domains in which they are applied. And notably, training agent models to predict specific actions presents a considerable challenge when trying to develop a single policy that can effectively learn multiple control systems across domains. Generally, the approach most modern works take when applying systems within specific domains is to start from a pretrained foundation model and then finetune a separate model for each specific domain. This fails to capture any commonalities between domains and results in a smaller total set of data used for training instead of leveraging each domain’s data.
7.3 Interactive agent for cross-modality and cross-reality
Developing AI agents that can successfully understand and perform tasks across different realities is an on-going challenge that has seen some recent success for image and scene generation (Huang et al., 2023a). In particular, it is challenging for agents to simultaneously understand real-world and virtual reality environments due to their visual dissimilarities and separate environment physics. Within the context of cross-reality, Sim to Real transfer is a particularly important problem when using simulation-trained policies for real-world data, which we discuss in the next section.
7.4 Simto Real Transfer
Techniques which enable models trained in simulation to be deployed in the real world. Embodied agents, especially one based on RL policies, are typically trained in simulated environments. These simulations do not fully replicate the characteristics of the real world (e.g., disturbances, light, gravity, and other physical properties). Due to this discrepancy between simulation and reality, models trained in simulation often struggle to perform well when applied in the real world. This issue is known as the “sim-to-real” problem. To solve this problem, several approaches can be taken:
• Domain randomization: domain randomization is a technique that trains a model while randomly varying parameters within a simulation environment (e.g., object appearance, sensor noise, and optical properties) in anticipation of the uncertainties and variations of the real world (Tobin et al., 2017). For instance, in the context of training a RL-based grasping skills, introducing randomness in the shapes of objects can lead to a policy capable of adapting to objects with somewhat different shapes (Saito et al., 2022).
• Domain adaptation: Domain adaptation, or domain transfer is a technique that bridges the gap between simulated and real-world domains by training models with a large number of simulated images and a smaller set of real-world images. In practical settings, unpaired image-to-image translation methods such as Cy cleGAN (Zhu et al., 2017b) are employed due to the difficulty in preparing paired images across domains. Several enhanced versions exist for reinforcement learning, including RL-CycleGAN (Rao et al., 2020), and for imitation learning, such as RetinaGAN (Ho et al., 2021).
• Improvement of simulation: Realistic simulation is a key for sim-to-real transfer. Part of this effort is achieved by a system identification techniques (Zhu et al., 2017c; Allevato et al., 2020), which aims to identify simulation parameters to mimic the real-world environments. Additionally, use of photorealistic simulators would be effective in image-based reinforcement learning (Martinez-Gonzalez et al., 2020; Müller et al., 2018; Shah et al., 2018; Sasabuchi et al., 2023). The sim-to-real transfer remains a central challenge in the study of Embodied Agents, as approaches keep evolving. Both theoretical and empirical research are essential to advance these technologies further.
8 Continuous and Self-improvement for Agent AI
Currently, foundation model based AI agents have the capacity to learn from multiple different data sources, which allow for more flexible sources for data for training. Two key consequences of this are (1) user and human-based interaction data can be used to further refine and improve the agent and (2) existing foundation models and model artifacts can be used to generate training data. We discuss each of these in more detail in the following sections, but we note that since current AI Agents are largely tied to existing pretrained foundation models, they generally do not learn from continuous interaction with their environments. We think this is an exciting future direction, and initial work by Bousmalis et al. has shown that self-improving agents for robotic control are able to continuous learn and improve through environmental interactions without supervision (Bousmalis et al., 2023). 8.1 Human-based Interaction Data The core idea behind using human-based interaction data is to leverage a large number of of agent-human interactions to train and improve future iterations of the agent. There are several strategies used to improve agents from human-agent interactions.
• Additional training data Perhaps the simplest usage of human-agent interactions is to use the interaction examples themselves as training data for a future iteration of the agent. This generally requires filtering strategies to differentiate successful agent examples from unsuccessful interaction examples. Filtering can be rules-based (e.g., reaching some desired end goal state), model-based (e.g., classifying successful vs unsuccessful interactions), or manually selected after a post hoc inspection and/or modification of the interaction examples.
• Human preference learning During interaction with the user, the agent system can prompt the user with several different model outputs and allow for the user to select the best output. This is commonly used by LLMs like ChatGPT and GPT-4, whereby users can select one output (out of several) that aligns best with their preferences.
• Safety training (red-teaming) Red-teaming within the context of Agent AI refers to having a dedicated team of adversaries (either human or computer) that seek to exploit and expose weaknesses and vulnerabilities within the Agent AI system. Although adversarial in nature, red-teaming is commonly used as a means for understanding how to improve AI safety measures and reduce the occurrence of harmful outputs. The core principle is to discover consistent methods for inducing unwanted agent outputs so that the model can be trained on data that explicitly corrects this behavior.
8.2 Foundation Model Generated Data
With the advent of powerful foundation model artifacts produced by academia and industry, there have been a variety of methods developed to extract and generate meaningful training data from these artifacts using a variety of prompting and data-pairing techniques.
• LLMInstruction-tuning Methods for generating instruction-following training data from LLMs have allowed for the finetuning of smaller, open-source models based on the outputs of larger proprietary LLMs (Wang et al., 2022b). For example, Alpaca (Taori et al., 2023) and Vicuna (Zheng et al., 2023) are LLMs based on the open-source LLaMA family (Touvron et al., 2023) that have been tuned on various outputs from ChatGPT and human participants. This method of instruction tuning can be viewed as a form of knowledge distillation, where the larger LLM serves as a teacher model to a smaller student model. Importantly, although LLM instruction-tuning has been shown to transfer the writing style and some instruction-following capabilities of the teacher model to the student model, significant gaps still exist between the factuality and capabilities of the teacher and student models (Gudibande et al., 2023).
• Vision-language pairs A number of recent works have sought to increase the number of diversity of pretraining data available to visual-language models by automatically generating captions and other text for visual content. For example, LLaVA (Liu et al., 2023c) uses 150,000 examples of instruction-following behavior from textual and visual inputs that are mainly LLM-generated. Other work has shown that using VLMs to re-caption images can improve the training data and subsequent quality of image generation models (Segalis et al., 2023). Within the realm of video understanding, using VLMs and LLMs to recaption videos has been shown to improve the performance and quality of subsequent VLMs trained on the recaptioned videos (Wang et al., 2023f; Zhao et al., 2022).
9 Agent Dataset and Leaderboard
To accelerate research in this domain, we propose two benchmarks respectively for multi-agent gaming and agentic visual language tasks. We will release two new datasets- “CuisineWorld” and “VideoAnalytica”- and a set of baseline models, encouraging participants to explore new models, systems, and submit their results on the test set of our leaderboard.
9.1 “CuisineWorld” Dataset for Multi-agent Gaming
CuisineWorld is a text-based game reminiscent of Overcooked! It offers a platform for AI-powered agents to cooperate and play in tandem. This dataset will test the collaboration efficiency of multi-agent systems, offering insights into how well LLMs and other systems can work together in dynamic scenarios. In particular, the dataset will focus on how well the agents understand goals, and how well the agents can coordinate among themselves. Two types of modes are supported in this dataset: a centralized dispatcher mode and a decentralized mode. Participants can choose a play mode and make a submission to our leaderboard.
9.1.1 Benchmark
For our competition, we will release a benchmark, the CuisineWorld benchmark, which includes a text interface that includes extendable task definition files, and an interface for multi-agent interaction, and human-machine interactions. Weintroduce the gaming interaction task in which the goal is to generate relevant, appropriate, multi-agent collaboration strategies that can maximize collaboration efficiency. We evaluate the collaboration efficiency with the proposed evaluation metric: CoS. The “CuisineWorld” dataset was collected by Microsoft, UCLA, and Stanford University. The goal of the competition is to explore how different, existing and novel, grounded-LLM and interactive techniques perform with this benchmark and establish strong baselines for the task of multi-agent gaming infrastructure. The dataset of CuisineWorld includes:- Aselection of well-defined multi-agent collaboration tasks.- An API system to facilitate agent interactions.- An automatic evaluation system. (The link for downloading the dataset will soon be made available and this article will be updated to include it here.)
9.1.2 Task
• Weprovide a dataset and related the benchmark, called Microsoft MindAgent and and correspondingly release a dataset “CuisineWorld” to the to the research community.
• Wewill provide benchmarks to evaluate and rank the submitted “MindAgent” algorithms. We will also provide baseline results generated using popular infrastructures. 9.1.3 Metrics and Judging The quality of multi-agent collaboration efficiency is determined by the new “cos” auto-metric (from MindAgent (Gong et al., 2023a)). The final rating of out metric is calculated as an average over the evaluated collaboration efficiency metrics of the multi-agent system on all tasks. Human evaluators will be asked to rate individual responses as well as provide subjective judgement of the engagement, breadth and an overall quality of the users’ interactions with the agents. 9.1.4 Evaluation
• Automated Evaluation. We plan to release a leaderboard, starting on the release date (TBA), registered participants will be asked to submit their results on the task associated with the dataset “CuisineWorld” (our publicly released dataset for the leaderboard). Submission of results will be closed on the end date (TBA). Each team will be required to submit their generated results on the testing set for automated evaluation of the “cos” metric.
• HumanEvaluation on our leaderboard. The leaderboard participants will need to provide a submission f ile generated by evaluation scripts locally. We will use the evalAI system to check the submission file and optionally rerun the code for top challenge contenders. Therefore, teams must also submit their code with a Readme file on how to run their code. Human evaluation will be performed by the organization team.
• Winner Announcement. We will make an announcement of the winners and post the final ratings of the submissions on our leaderboard.
9.2 Audio-Video-Language Pre-training Dataset.
We introduce VideoAnalytica: a new benchmark for analytical video demonstration comprehension. VideoAnalytica focuses on leveraging video demonstrations as aids to better understand complex, high-level reasoning embedded within long-formed instructional videos. The objective is to evaluate the cognitive reasoning abilities of video language models, pushing them beyond mere recognition tasks and basic comprehension, towards a more sophisticated and nuanced understanding of videos. Crucially, VideoAnalytica emphasizes the integration of multiple modalities, such as audio, video, and language, as well as the ability of models to apply domain-specific knowledge, to contextualize and interpret the information presented in the videos. Specifically, VideoAnalytica involves two primary tasks: 1. Video Text Retrieval: This task involves accurately retrieving relevant text from the instructional videos. The challenge lies in distinguishing between relevant and irrelevant information, thus requiring a deep understanding of the video content, and analysis of the demonstration to retrieve the correct query. To further increase the complexity of these tasks, we introduce hard negatives into our datasets generated by large language models. We run human validation on the generated negatives and remove instances that make the task invalid and unfair (e.g. negatives being valid). 2. Video Assisted Informative Question Answering: This task requires the model to answer questions based on the information extracted from the videos. The focus is on complex questions that require analytical reasoning and a thorough comprehension of the video demonstration. To facilitate the development of an audio-video-language agent for analytical video understanding, we introduce a benchmark leaderboard for the two tasks from VideoAnalytica.
• The leaderboard participants will need to submit their solutions for evaluation. The evaluation will be based on the model’s performance on the two tasks, and the results will be displayed on the leaderboard. Participants are required to submit their code, along with a detailed explanation of their approach and methodology.
• Ethical considerations: The leaderboard focuses on understanding and interpreting video content, which could potentially be used in surveillance or other privacy-invasive applications. Therefore, it’s crucial to consider the ethical implications and potential misuse of the technology. We encourage participants to consider these aspects in their submissions and promote the ethical use of AI.
10 Broader Impact Statement
This article and our associated forum (https://multimodalagentai.github.io) aim to be a catalyst for innovative research, fostering collaborations that will drive the next wave of AI applications. By focusing on multimodal agents, we emphasize the future direction of human-AI interactions, leader-board, and solutions. We detail three ways in which we make significant contributions to the broader community.
Firstly, we hope our forum grounds AI researchers to develop solutions motivated by real-world problems in gaming, robotics, healthcare, and long-video understanding. Specifically, the development of multimodal agents in gaming could lead to more immersive and personalized gaming experiences, thereby transforming the gaming industry. In robotics, the development of adaptive robotic systems could revolutionize industries ranging from manufacturing to agriculture, potentially addressing labor shortages and improving efficiency. In healthcare, the use of LLMs and VLMs as diagnostic agents or patient care assistants could lead to more accurate diagnoses, improved patient care, and increased accessibility to medical services, particularly in underserved areas. Furthermore, the ability of these models to interpret long-form videos could have far-reaching applications, from enhancing online learning to improving technical support services. In general, the topics covered in our forum will have significant downstream effects on a wide range of industries and humans across the world.
Secondly, we hope our forum stands as a valuable resource for AI practitioners and researchers alike, serving as a platform to explore and deeply comprehend the diverse and complex leader-board that come with implementing AI agents across a wide variety of environments and situations. This exploration includes, for instance, understanding the specific limitations and potential hazards linked to Agentic AI systems when they are developed for specialized sectors such as healthcare diagnostics. In this domain, issues like dangerous hallucinations in AI behavior can pose significant risks, highlighting the critical need for meticulous design and testing. However, these specific leader-board may not be equally relevant or noticeable when considering AI agents crafted for the gaming industry. In such recreational fields, developers might instead prioritize tackling different hurdles, such as the need for AI to perform more open-ended generation and exhibit creativity, adapting dynamically to unpredictable gameplay scenarios and player interactions. By attending the forum, participants will gain insights into how these varied environments dictate the focus and direction of AI development, and how best to tailor AI solutions to meet these distinct needs and overcome the pertinent leader-board.
Thirdly, the various elements of our event, including the expert presentations, informative posters, and notably the winners of our two leader-board, are set to offer a substantive yet succinct overview of the latest and significant trends, research directions, and innovative concepts in the realm of multimodal agents. These presentations will encapsulate pivotal findings and developments, shining a light on new systems, ideas, and technologies in the field of mulitmodal agent AI. This assortment of knowledge is not only beneficial for the attendees of our forum, who are looking to deepen their understanding and expertise in this domain, but it also serves as a dynamic and rich resource board. Those visiting our forum’s website can tap into this reservoir of information to discover and understand the cutting-edge advancements and creative ideas steering the future of multimodal agent AI. We strive to serve as a useful knowledge base for both newcomers and veterans in the field. By engaging with these resources, we hope participants and online visitors alike can remain informed of the transformative changes and novel approaches that are shaping the exciting landscape surrounding multimodal agent AI.
11 Ethical Considerations
Multimodal Agent AI systems have many applications. In addition to interactive AI, grounded multimodal models could help drive content generation for bots and AI agents, and assist in productivity applications, helping to re-play, paraphrase, action prediction or synthesize 3D or 2D scenario. Fundamental advances in agent AI help contribute towards these goals and many would benefit from a greater understanding of how to model embodied and empathetic in a simulate reality or a real world. Arguably many of these applications could have positive benefits. However, this technology could also be used by bad actors. Agent AI systems that generate content can be used to manipulate or deceive people. Therefore, it is very important that this technology is developed in accordance with responsible AI guidelines. For example, explicitly communicating to users that content is generated by an AI system and providing the user with controls in order to customize such a system. It is possible the Agent AI could be used to develop new methods to detect manipulative content- partly because it is rich with hallucination performance of large foundation model- and thus help address another real world problem. For examples, 1) in health topic, ethical deployment of LLM and VLM agents, especially in sensitive domains like healthcare, is paramount. AI agents trained on biased data could potentially worsen health disparities by providing inaccurate diagnoses for underrepresented groups. Moreover, the handling of sensitive patient data by AI agents raises significant privacy and confidentiality concerns. 2) In the gaming industry, AI agents could transform the role of developers, shifting their focus from scripting non-player characters to refining agent learning processes. Similarly, adaptive robotic systems could redefine manufacturing roles, necessitating new skill sets rather than replacing human workers. Navigating these transitions responsibly is vital to minimize potential socio-economic disruptions. Furthermore, the agent AI focuses on learning collaboration policy in simulation and there is some risk if directly apply ing the policy to the real world due to the distribution shift. Robust testing and continual safety monitoring mechanisms should be put in place to minimize risks of unpredictable behaviors in real-world scenarios. Our “VideoAnalytica” dataset is collected from the Internet and considering which is not a fully representative source, so we already go through-ed the ethical review and legal process from both Microsoft and University Washington. Be that as it may, we also need to understand biases that might exist in this corpus. Data distributions can be characterized in many ways. In this workshop, we have captured how the agent level distribution in our dataset is different from other existing datasets. However, there is much more than could be included in a single dataset or workshop. We would argue that there is a need for more approaches or discussion linked to real tasks or topics and that by making these data or system available.
We will dedicate a segment of our project to discussing these ethical issues, exploring potential mitigation strategies, and deploying a responsible multi-modal AI agent. We hope to help more researchers answer these questions together via this paper.
12 Diversity Statement
By examining the adaptability of AI agent models in various domains, we inherently embrace a diversity of leader-board, perspectives, and solutions. In this vein, our project aims to build a diverse community by exploring the wide array of subjects in multimodal and agentic AI.
With these principles in mind, this project focuses on advanced multimodal systems that interact effectively within both physical and virtual environments and facilitate effective interaction with humans. As such, we intend to engage a broad range of experts and practitioners across a wide-range of technical specialities, cultures, countries, and scholarly f ields to discuss important topics, including but not limited to:
• Application of foundation models: the development of agents with integrated modalities (audio, image, text, sensor inputs), aiming to enhance their recognition and response capabilities for a wide variety of applications. • General-purpose end-to-end systems: the development of end-to-end models that are trained with large-scale data, seeking to create versatile and adaptable AI solutions. • Methodologies for grounding modalities: integrating information across various modalities, enhancing the coherence and efficacy of data processing. • Intuitive human interface: the development of effective and meaningful interaction between humans and agents. • Taming LLM/VLMs: exploring new approaches to address common issues in large-scale models, such as hallucinations and biases in their outputs.
We aspire to broaden our collective understanding of the potential and limitations of agentic AI by leveraging our unique and diverse perspectives. We strongly believe that this approach will not only enrich individual perspectives, but will also enhance the community’s collective knowledge and promote a holistic view that is more inclusive of the wide-ranging leader-board faced by multimodal AI agents.
[6]Blackburn, S., 1995, “Practical Tortoise Raising”, in Mind 104.
[7]Brown, D. G., 1954, “What the Tortoise Taught Us”, in Mind 63.
[8]Brunero, J., 2005, “Instrumental Rationality and Carroll’s Tortoise”, in Ethical Theory Moral Practice 8.
[9]Carroll, L., 1895, “What the Tortoise Said to Achilles”, in Mind 4.
[10]Fumerton, R., 2015, “What the Internalist Should Say to the Tortoise”, in Episteme 12.
[11]Husserl, E., 1968, Ph?nomenologische Psychologie: Vorlesungen Sommersemester 1925, W. Biemel (hrsg.), Den Haag: Martinus Nijhoff.
1976, Ideen zu einer reinen Ph?nomenologie und ph?nomenologische Philosophie. Erstes Buch, K. Schuhmann (hrsg.), Den Haag: Martinus Nijhoff.
1991, Ideen zu einer reinen Ph?nomenologie und ph?nomenologische Philosophie. Zweites Buch, M. Biemel (hrsg.), Dordrecht: Kluwer Academic Publishers.
[12]Irvine, A. D., 1996, “Philosophy of Logic”, in S. G. Shanker(ed.),Routledge History of Philosophy, Volume IX, Philosophy of Science Logic and Mathematics in the Twentieth Century, London: Routledge.
[13]Murata, N., 2019, “How is Time Constituted in Consciousness? Theories of Apprehension in Husserl’s Phenomenology of Time”, in N. de Warren and S. Taguchi (eds.), New Phenomenological Studies in Japan, Cham: Springer.
[14]Railton, P., 1997, “On the Hypothetical and Non-hypothetical in Reasoning about Belief and Action”, in G. Cullity and B. Gaut (eds.), Ethics and Practical Reason, Oxford: Clarendon Press.
[15]Rees, W. J., 1951, “What Achilles Said to the Tortoise”, in Mind 60.
[16]Russell, B., 1903, The Principles of Mathematics, Cambridge: Cambridge University Press.
[17]Ryle, G., 2009, “If, So, and Because”, in Collected Papers Volume 2: Collected Essays 1929-1968, London: Routledge.
[18]Schueler, G. F., 1995, “Why ‘Oughts’ are not Facts”, in Mind 104.
[19]Smiley, T., 1995, “A Tale of Two Tortoises”, in Mind 104.
[20]Stroud, B., 1979, “Inference, Belief, and Understanding”, in Mind 88.
[21]Tieszen, R., 2011, After G?del, New York: Oxford University Press.
[22]Thomson, J. F., 1960, “What Achilles Should Have Said to the Tortoise”, in Ratio 3.
[23]Wieland, J. W., 2013, “What Carroll’s Tortoise Actually Proves”, in Ethical Theory Moral Practice 16.
蒯因直接采纳塔尔斯基的真定义,并认为真谓词的功能就是去引号,即起着从提及语句到使用语句的转换的作用。霍里奇的最小论实际上也是由类似(T)这样的实例所构成,区别在于使用命题而不是语句作为真值载体。按照蒯因和霍里奇的理解,塔尔斯基的真理论无疑是一个收缩论。戴维森尽管没有认同这一点,但却明确否认塔尔斯基和亚里士多德是符合论者,理由是在他们的真定义或真概念中缺失了符合论所需要的事实或事态概念。[8],第268页与之相反,谢尔则认为塔尔斯基和亚里士多德的真概念都是符合论的概念,理由在于其背后的根本观点是:一个语句为真不但跟该语句所说的东西有关,也跟世界中的事物是如何的(how things are)有关。[25],第135页
②“Deflation/inflation”的物理意义是“放气、缩小/充气、膨胀”,经济学意义是“通货紧缩/通货膨胀”。真理论中的“deflate/inflate”是借用物理上的隐喻,例如“deflate the overinflated balloons offered by substantivists”。[6],第4页故相应地,笔者将“deflationism/inflationism”译为“收缩论/膨胀论”。
[1]W.Alston,2002,”Truth:Concept and property”,in R.Schantz(ed.),What is Truth,pp.11-26,New York:de Gruyter.
[2]J.Asay,2013,The Primitivist Theory of Truth,Cambridge:Cambridge University Press.
[3]P.Boghossian,1990,”The status of content”,Philosophical Review,99(2):157-184.
[4]J.Cleve,1996,”Minimal truth is realist truth”,Philosophy and Phenomenological Research,56(4):869-875.
[5]N.Damnjanovic,2010,”New wave deflationism”,in C.D.Wright and N.J.L.L.Pedersen(eds.),New Waves in Truth,pp.45-58,New York:Palgrave Macmillan.
[6]M.David,1994,Correspondence and Disquotation:An Essay on the Nature of Truth,Oxford:Oxford University Press.
[7]D.Davidson,1990,”The structure and content of truth”,Journal of Philosophy,87(6):279-328.
[8]D.Davidson,1996,”The folly of trying to define truth”,Journal of Philosophy,93(6):263-278.
[9]M.Dummett,1991,The Logical Basis of Metaphysics,Cambridge MA:Harvard University Press.
[10]D.Edwards,2013,”Truth as a substantive property”,Australasian Journal of Philosophy,91(2):279-294.
[11]H.Field,1972,”Tarski’s theory of truth”,Journal of Philosophy,69(13):347-375.
[12]H.Field,1986,”The deflationary conception of truth”,in G.Macdonald and C.Wright(eds.),Fact,Science and Morality:Essays on A.J Ayer s Language,Truth and Logic,pp.55-117,Oxford:Blackwell.
[13]H.Field,1994,”Deflationist views of meaning and content”,Mind,103(411):249-285.
[14]H.Field,1994,”Disquotational truth and factually defective discourse”,Philosophical Review,103(3):405-452.
[15]H.Glock,2008,What is Analytic Philosophy?,Cambridge:Cambridge University Press.
[16]P.Horwich,1982,”Three forms of realism”,Synthese,51(2):181-201.
[17]P.Horwich,1998,Truth(2nd edition),Oxford:Oxford University Press.
[18]P.Horwich,2018,”Is truth a normative concept?”,Synthese,195(3):1127-1138.
[19]R.Kirkham,1992,Theories of Truth:A Critical Introduction,Cambridge MA:MIT Press.
[20]W.Künne,2003,Conceptions of Truth,Oxford:Oxford University Press.
[21]D.Lewis,1983,”New work for a theory of universals”,Australasian Journal of Philosophy,61(4):343-377.
[22]M.Lynch,2009,Truth as One and Many,Oxford:Oxford University Press.
[23]D.Patterson,2003,”What is a correspondence theory of truth?”,Synthese,137(3):421-444.
[24]W.V.Quine,1970,Philosophy of Logic,Englewood Cliffs:Prentice-Hall.
[25]G.Sher,1998,”On the possibility of a substantive theory of truth”,Synthese,117(1):133-172.
[26]G.Sher,2004,”In search of a substantive theory of truth”,Journal of Philosophy,101(1):5-36.
代谢性疾病作为一类由机体代谢紊乱引起的疾病,其中代谢性骨病(Metabolic bone disease)是古代人类遗骸中较为常见的一种类型,具体指导致正常骨形成、吸收或矿化发生系统性改变的疾病或疾病组合,多与营养不良和激素失调有关。根据现代医学研究成果,代谢性骨病早期发病特征不典型,中后期临床表现复杂,常见生长障碍、骨关节病、骨骼畸形。中后期代谢性骨病在骨骼上具有显著的体质特征表现,易于在考古材料中观察识别,因此,这类疾病在古代人类遗骸中观察到的比率较大,主要涉及骨质疏松症、氟骨症、佝偻病和坏血病等类型,目前针对代谢性骨病的多学科研究也在持续开展。
骨质疏松症的致病因素较为复杂,通常认为与年龄、性别、饮食及营养状况有关,特定生活方式也可能诱发OP,可以此为线索探究古代人群社会分工和生活方式差异。例如,M. E. Zaki等学者对公元前2687年至前2191年古埃及人骨骼遗骸进行骨矿物质密度(Bone Mineral Density,BMD)检测,这些骨骼来自两个不同社会阶层:高级官员和工人群体。结果显示BMD值与年龄、性别和社会身份存在关联,有关年龄和社会身份的差异具体表现为老年群体的BMD值较年轻群体明显下降,且男性工人的骨质疏松症发病率高于男性高级官员,而女性高级官员的发病率则高于女性工人。研究者认为不同群体的致病原因存在差异,推测男性工人骨质疏松症发病率较高可能与营养不足和过重的工作量有关,而女性高级官员的久坐生活方式则是潜在致病因素之一。此外,女性骨质疏松症的平均发病时间早于男性且发病频率更高,这种现象可能与女性更年期的荷尔蒙变化有关。国内相关学者针对古代人类遗骸的骨质疏松症开展过诸多方面的研究,如郑晓瑛对甘肃酒泉干骨崖墓地出土的青铜时代人骨进行了X-光病理鉴定,不仅确认了氟骨症的发病证据以及骨包虫病和骨肿瘤病的可能性,还发现样本骨质疏松症发病年龄呈现出低于现代人发病年龄的倾向,古代特殊的生存环境与生活状况可能导致了发病年龄的差异。王明辉比较了贾湖遗址和西坡墓地出土人骨骨质疏松症的发病率,指出西坡农业人群的高发病率除了可能存在的流失钙质的疾病外,应与人群间饮食和营养状况的差距有关,早期生业模式的转变可能提升了骨质疏松发病率。
目前古代人类遗骸样本中坏血病证据跨越了数千年,几乎遍布全世界,较早的病例来自公元前3000年左右的德国、希腊和约旦地区。早期报告的坏血病病例主要集中在成年人群体,随着诊断方法的完善,如今绝大多数病例都发现于青少年群体。通过对不同地区古代人群的研究,相关学者对古代坏血病的病理特征与致病因素有了更深入的认识。Haagen D. Klaus在南美洲出土的青少年骨骼表面发现颅外血管压痕,这表明古代青少年患者可能会出现积血症状,对古代坏血病的体质特征做出了补充。Anne Marie E. Snoddy等人观察到来自智利阿塔卡马沙漠的距今约3400年的四具新生儿遗骸呈现出非特异性的骨骼畸变,其中一个新生儿与同一地区出土一位罹患坏血病的成年女性存在血缘关系,这显示出沙漠地区农业转型时期的资源短缺可能对产妇与胎儿的健康都造成了负面影响,但受限于诊断技术,将目前的坏血病诊断标准应用于新生儿遗骸仍面临诸多挑战。Chryssi Bourbou通过研究11—12世纪希腊的青少年遗骸,成功发现青少年坏血病的证据,不仅丰富了该地区这一疾病的历史病例,还提出青少年坏血病的发生可能与断奶后摄入固体食物的种类与品质有关。综合多项研究可见,坏血病患病概率很可能与生活方式、资源获取及文化因素有关。
古代疾病研究始终面临证据碎片化、疾病表现复杂性等诸多挑战,单一学科的研究方法往往难以全面且准确地揭示疾病的本质特征与演变规律。在此背景下,多学科研究范式逐渐成为古代疾病研究的核心路径。该范式整合传统体质人类学、古分子生物学、考古学、历史学等多学科的理论与技术手段,通过学科交叉融合,不仅能够从骨骼病变等体质特征中获取直观信息,还能借助古 DNA 分析、蛋白质检测等前沿技术深入探究疾病在分子层面的发作机制、演化轨迹以及传播路径。无论解析代谢性疾病的致病原理,还是诊断骨骼特异性感染并构建疾病时空框架,多学科研究范式均展现出显著优势,有助于系统理解古代疾病的出现、发展及其对人类健康和社会文化的影响。
佩吉特骨病(Paget disease of bone,PDB)又称变形性骨炎,是一种慢性骨代谢疾病,其病理机制在细胞层面表现为破骨细胞增大、增多,同时伴随成骨细胞增加且矿化不良,致使骨形成加快6至7倍。这种代谢异常会导致新骨混乱,表现为骨骼外观增大、外表坚硬、但骨质量较差的病症,多累及盆骨、颅骨、长骨等骨骼,并且极易引发骨折、骨肉瘤或关节炎等并发症。PDB的病因尚不明晰,早期认为佩吉特骨病或与人畜共患的传染病有关,现代遗传学研究表明PDB有家族遗传倾向,患者的CSF1、OPTN和TNFRSF11A三种基因更易出现缺陷。
软骨发育不全已经存在了数千年,根据历史记载和艺术作品的相关描绘,古代埃及应有不少相关病例,目前较早的骨骼证据来自新石器时代中期的法国。分子生物学研究进一步提升了对软骨发育不全的认识,Lucas L. Boer等学者在一例180年前软骨发育不全骨骼样本中检测到了FGFR3基因编码的杂合子G1138A变异,以历史证据有力证实了该病症是由基因FGFR3的致病性错义突变引起。对于体质特征不典型或不明确的古代人类骨骼遗存,通过古分子研究证实其携带FGFR3基因特定的致病性变异可以辅助诊断。
肿瘤可分为良性和恶性两大类。良性肿瘤往往在原发病灶独自生长,仅在局部扩散;恶性肿瘤是指原发生长物向身体其他器官无限制地局部扩散,在肿瘤发生部位可能还伴有自发性骨折。国内外古代样本中可由体质特征判断的的良性肿瘤包括骨瘤、骨样骨瘤、骨血管瘤等。骨血管瘤(Hemangiomas)为原发于骨血管的良性肿瘤,是一种掺杂于骨小梁之间的呈瘤样增生的血管组织,好发于扁骨,如脊柱、颅骨、颌骨,长骨少见,分为海绵型和毛细血管型。在考古样本中可辨别血管瘤的皮质破坏属于疾病晚期特征,多发生于老年个体,因此考古样本中呈现的骨血管瘤患病率远低于现代临床样本,如Joseph E. Molto等研究者在一具埃及地区出土的罗马时期老年女性骨骼上诊断出晚期脊柱血管瘤(Vertebral hemangiomas,VHs)。恶性肿瘤包括骨肉瘤、多发性骨髓瘤、骨转移癌等。骨转移癌(Metastatic Carcinoma)是指原发于某些器官的恶性肿瘤通过血液循环、淋巴系统或脑脊液转移到骨骼的继发性恶性肿瘤,通过再转移或直接浸润到骨骼造成骨破坏。例如张群等对宁夏石砚子墓地出土的东汉时期颅骨缺损个体开展体质观察,血管压迹的增粗和加深显示出滋养肿瘤的异常血管形态,推测个体颅骨所见的大面积骨质破坏应由颅骨转移癌所导致。
古分子研究的突破性贡献,则在于其打破了病原体演化与人类历史的时空壁垒,首次从分子层面证实疾病是塑造人群迁徙与文明格局的核心力量。对鼠疫耶尔森氏菌基因组的重建,揭示了新石器时代末期该病原体通过欧亚贸易网络扩散的轨迹,印证了物质交流与疾病传播同步发生的历史逻辑,填补了史前流行病如何影响人口结构与文化更替的认知空白;对疟疾、乙型肝炎病毒的古 DNA 分析,不仅追溯了跨大西洋奴隶贸易、欧亚人群迁徙中的病原体传播路径,更发现了人类与病原体的基因共演化证据(如抗疟基因的自然选择),使我们意识到健康适应本身就是人类演化的重要驱动力—这种认知突破,让疾病史从边缘学科话题跃升为理解人类历史进程的核心维度之一。而多学科研究范式的成熟,则标志着古代疾病研究进入系统阐释文明互动的新阶段。通过整合体质人类学的病理观察、古分子生物学的基因证据、稳定同位素的饮食分析与考古学的文化背景等,我们得以搭建疾病、社会与文明的完整认知链条,例如对软骨发育不全的研究,不仅通过FGFR3基因变异确诊疾病,更结合墓葬位置、社群布局推测古代社会对身体缺陷的接纳程度;对肿瘤的多学科分析,通过组织学、古DNA与同位素结合,揭示了饮食结构(如高肉食摄入)与基因变异(如 K-ras 突变)的关联,为理解古代生活方式如何影响疾病发生提供了立体视角。这种跨学科融合,彻底超越了单一学科的局限,使疾病如何参与文明演进的深层探讨成为可能。
[6]Blackburn, S., 1995, “Practical Tortoise Raising”, in Mind 104.
[7]Brown, D. G., 1954, “What the Tortoise Taught Us”, in Mind 63.
[8]Brunero, J., 2005, “Instrumental Rationality and Carroll’s Tortoise”, in Ethical Theory Moral Practice 8.
[9]Carroll, L., 1895, “What the Tortoise Said to Achilles”, in Mind 4.
[10]Fumerton, R., 2015, “What the Internalist Should Say to the Tortoise”, in Episteme 12.
[11]Husserl, E., 1968, Ph?nomenologische Psychologie: Vorlesungen Sommersemester 1925, W. Biemel (hrsg.), Den Haag: Martinus Nijhoff.
1976, Ideen zu einer reinen Ph?nomenologie und ph?nomenologische Philosophie. Erstes Buch, K. Schuhmann (hrsg.), Den Haag: Martinus Nijhoff.
1991, Ideen zu einer reinen Ph?nomenologie und ph?nomenologische Philosophie. Zweites Buch, M. Biemel (hrsg.), Dordrecht: Kluwer Academic Publishers.
[12]Irvine, A. D., 1996, “Philosophy of Logic”, in S. G. Shanker(ed.),Routledge History of Philosophy, Volume IX, Philosophy of Science Logic and Mathematics in the Twentieth Century, London: Routledge.
[13]Murata, N., 2019, “How is Time Constituted in Consciousness? Theories of Apprehension in Husserl’s Phenomenology of Time”, in N. de Warren and S. Taguchi (eds.), New Phenomenological Studies in Japan, Cham: Springer.
[14]Railton, P., 1997, “On the Hypothetical and Non-hypothetical in Reasoning about Belief and Action”, in G. Cullity and B. Gaut (eds.), Ethics and Practical Reason, Oxford: Clarendon Press.
[15]Rees, W. J., 1951, “What Achilles Said to the Tortoise”, in Mind 60.
[16]Russell, B., 1903, The Principles of Mathematics, Cambridge: Cambridge University Press.
[17]Ryle, G., 2009, “If, So, and Because”, in Collected Papers Volume 2: Collected Essays 1929-1968, London: Routledge.
[18]Schueler, G. F., 1995, “Why ‘Oughts’ are not Facts”, in Mind 104.
[19]Smiley, T., 1995, “A Tale of Two Tortoises”, in Mind 104.
[20]Stroud, B., 1979, “Inference, Belief, and Understanding”, in Mind 88.
[21]Tieszen, R., 2011, After G?del, New York: Oxford University Press.
[22]Thomson, J. F., 1960, “What Achilles Should Have Said to the Tortoise”, in Ratio 3.
[23]Wieland, J. W., 2013, “What Carroll’s Tortoise Actually Proves”, in Ethical Theory Moral Practice 16.
突变理论的模型的性质,最好用例子来说明,我们从研究狗的进攻模型开始。洛仑兹(Konrad Z. Lorenz)曾指出,进攻行动受两个互相矛盾的倾向所制约:发怒和恐惧。他还指出,对于狗来说,这两种因素在某种程度上可以测量出来。一只狗的发怒和张嘴、露齿程度有关,其恐惧程度则可从它的耳朵向后拉平多少反映出来。使用面部表情作为狗的情绪状态的指标,我们可望弄清狗的行为的变化是如何因情绪变化而变化的。
增加控制空间和行为空间的维数,可以构造出无限的突变序列。俄国数学家阿诺尔德(V. I. Arnold)已经至少对25维进行了分类。但在现实世界的现象模型中,是上面所描述的七种可能最为重要,因为它们具有不超过四维的控制空间。由空间位置和时间所决定的各种过程的特殊同类性,不能多于四维的控制空间,因为我们的世界只有空间三维和时间一维。
我们仅考虑那些在训练语料库中实际出现或出现频率足够高的连续词组合。当出现一个训练语料库中未曾见过的 n 词新组合时会发生什么?我们不应给这种情况分配零概率,因为这类新组合很可能出现,且上下文窗口越大时出现频率会更高。一个简单的解决方案是参考更小上下文窗口预测的概率,如回退三元模型(Katz, 1987)或平滑(插值)三元模型(Jelinek and Mercer, 1980)所做的那样。那么在这类模型中,如何实现从训练语料库中观察到的词序列到新词序列的基本泛化?理解这一机制的方式是设想一个与这些插值或回退 n 元模型对应的生成模型。本质上,新词序列是通过”粘合”训练数据中高频出现的、长度为 1、2…至 n 的极短重叠片段而生成的。 获取下一个词块概率的规则隐含在回退或插值 n-gram 算法的具体实现中。研究者通常采用 n=3(即三元模型)并获得了最先进的成果,但 Goodman(2001 年)的研究表明,结合多种技巧可带来显著提升。显然,待预测词之前的序列信息远不止前几个词的身份标识。这种方法至少存在两个亟待改进的特征,而我们将在本文中重点探讨这些获得最先进成果的改进方向。首先,现有方法未能考虑超过 1-2 个词之外的上下文;其次,它忽略了词语之间的”相似性”。例如,当训练语料中出现过”猫在卧室里行走”这样的句子时,应当能帮助我们泛化推断出”狗在房间里奔跑”也具有相近的概率,这仅仅因为”狗”与”猫”(以及”the”与”a”、”room”与”bedroom”等)具有相似的语义和语法角色。
若并行计算机由多个 CPU 构成网络架构,由于参数交换量庞大(最大规模网络下参数规模接近 100MB),频繁在处理器间传输全部参数将超出本地网络带宽的承受能力。为此我们采用参数并行化策略,重点针对输出单元参数进行划分——这正是我们架构中计算最密集的环节。每个 CPU 负责计算部分输出单元的非归一化概率,并更新对应输出单元的权重参数。该方案实现了通信开销极低的高效并行随机梯度上升,各 CPU 仅需交换两类关键数据:(1)输出层 softmax 的归一化因子;(2)隐藏层(下文记作 a )与词特征层(下文记作 x )的梯度信息。 所有 CPU 都会重复计算输出单元激活前的运算,包括词特征选择、隐藏层激活 的计算,以及相应的反向传播和参数更新步骤。不过对于我们的网络架构而言,这些计算量仅占总计算量的极小部分。
以美联社(AP)新闻数据实验采用的架构为例:词汇表大小 |V|=17964 ,隐藏单元数 h=60 ,模型阶数 n=6 ,词特征维度 m=100 。处理单个训练样本所需的总运算量约为 |V|(1+nm+h)+h(1+nm)+nm(其中各项分别对应输出单元、隐藏单元和词模型阶数 n=6 的特征单元计算)。在此例中,输出单元加权求和计算量约占整体计算量的比例约为99.7%。这个计算是近似值,因为不同操作实际消耗的 CPU 时间存在差异,但它表明并行化输出单元的计算通常具有优势。对于本文追求的并行化程度(即几十个处理器)而言,所有 CPU 重复执行极小部分计算并不会显著影响总计算时间。如果隐藏单元数量庞大,并行化其计算也将变得有利可图,但我们在实验中未对该方法进行深入研究。
该策略的实施是在一个由 1.2 GHz 主频 Athlon 处理器(32 台×2 CPU)组成的集群上完成的,这些处理器通过 Myrinet 网络(一种低延迟千兆位局域网)连接,并使用 MPI(消息传递接口)库(Dongarra 等人,1995 年)进行并行化处理。以下简要描述针对单个样本 (wt-n+1,…,wt) 的并行算法,该算法由集群中 M 个处理器中的第 i 个 CPU 并行执行。CPUi (i 的范围从 0 到M-1 )负责从编号 starti=i×[|V|/M]开始的输出单元块,该块的长度为 min([|V|/M,|V|]- starti) 。
上述神经网络的一个变体可被解释为遵循Hinton’s近期关于专家乘积(Hinton, 2000)研究的能量最小化模型。在前文描述的神经网络中,分布式词特征仅用于”输入”词而不用于”输出”词(下一个词)。此外,输出层扩展了极大量参数(占大多数),且未利用输出词之间的语义或句法相似性。此处描述的变体中,输出词同样由其特征向量表示。该网络接收单词子序列(映射为其特征向量)作为输入,并输出能量函数 E ——当单词构成可能子序列时 值较低,不可能时值较高。例如,该网络输出的”energy”函数为:
其中 b 是偏置向量(对应无条件概率), d 是隐藏单元偏置向量,v 是输出权重向量,H 是隐藏层权重矩阵。与先前模型不同,此处输入词和输出词共同构成 x :
能量函数 E(wt-n+1,…,wt) 可解释为 (wt-n+1,…,wt) 联合出现的非归一化对数概率。要获得条件概率 ,只需^P(wt|wt-1t-n+1)(尽管计算代价较高)对 w 的可能值进行归一化处理,具体如下:
耶路撒冷东北25千米的Tell el Sultan是《圣经》上反复出现的一个地名,1870年代,考古发掘证实,这里实际上属于《圣经》中提及次数更多的古城耶利哥(Jericho)的一部分。后来,这一带陆续出土的古人类定居的遗址逐渐超过20处,时间大多超过4 000年。1952-1958年,英国的女考古学家凯瑟琳·凯尼恩(Kathleen Kenyon,1906-1978)主持对新的土层进行系统挖掘,彻底改变了人们对历史的看法:耶利哥的早期人类遗迹超过一万年。
1968年,人们在叙利亚境内的幼发拉底河上建造塔巴水坝(Tabqa Dam)时,挖掘出一个人类居住了近4 000年(1.1万——0.75万年前)的遗址——阿布·胡列伊拉(Tell Abu Hureyra)。这是一个从狩猎采集生活形态向农业种植形态过渡的遗址,这里的生活者也因此被称为世界上最早的农民。在这个遗址,从土壤和动物鱼骨等物质中成功分离出712个种子样本,最终查明属于150类以上食用植物的500多种植物种子。这个叙利亚遗址再现了1.1万年前的人类采集狩猎生活方式,和大约一万年前开始的初步的农业种植生活的轮廓。
卡拉卡山区丰富的可食用植物种子和种植技术,沿着黎巴嫩——以色列——叙利亚——伊拉克,一直传播到地中海沿岸。其中最著名的遗迹包括耶利哥(Jericho)、叙利亚的阿布·胡列伊拉遗址(Tell Abu Hureyra)、土耳其的加泰土丘(Catal Huyuk)。(加泰土丘挖掘时间在1950-1990年,新石器时代的14层遗迹厚度达15米,时间为8 850年前,当时已经进入所谓的金石混用时代(中东学者对青铜时代的称呼),艺术和宗教的文物极其丰富)
这项研究没有取得什么结果。血型和细胞表面蛋白质标记无法确认人的血统世系,也无法落实迁移路线。这是当时的研究技术的限制。但是,斯福扎发现农业并非单纯的文化现象,而是伴随着人口的快速增长,这股风潮从欧洲的东南部向西北部扩散,后来被称为“前进的浪潮”(Wave of Advance)。这种“前进的浪潮”被很多人接受了,但是斯福扎本人并不接受这种观念,因为人们还没有搞清楚欧洲的基因库的起源。
美国国家健康研究院(National Institutes of Health)的迪尔德丽·乔伊(Deirdre Joy)和她的同事发现,疟原虫在5万年前开始多样化,这个时间恰好是人类走出非洲的时期,暗示人类带着疟原虫前往世界各地。乔伊还发现了其他证据,一万年前,疟原虫开始大规模的多样化,这个时间正是新石器革命的农业起源的时间。
G6PD是细胞里的一种酶,可以把葡萄糖转化成一种亚细胞能量包(subcellular energy packet),这种亚细胞能量包名为NADPH,是人类细胞能量活力的来源。我们吃下的谷物——碳水化合物又称多糖类,被转化为单糖(葡萄糖)后,最终变成我们细胞里的三种能量:NADPH、NADH和ATP。所以G6PD极其重要。
威廉·布莱船长(William Bligh,1754-1817)的故事《叛舰喋血记》(The Mutiny of the Bounty)曾5次被搬上银幕。1789年,他率领的“邦蒂号”(Bounty)经过6个月航行抵达塔希提。他一路上都严苛地虐待水手,抵达塔希提后,他强令水手不许寻找当地女人以免传染性病。
第一个指出这种风险的学者是日本裔美国生物学家大野乾(Susumu Ohno,1928-2000),他在1970年所著的《基因重复的进化》(Evolution by GeneDuplication)一书中提出:重复基因时,随心所欲地草率选择,会导致“快速进化”的变异,必须保留备份。他创造出“垃圾DNA”(junk DNA)一词,用以描述基因组里的很多功能不详的DNA。这种垃圾是重复基因的必然宿命,也许毫无意义,也许后果致命。
世界自然遗产大烟山国家公园(Great Smoky Mountains National Park)是美国旅游人数最多的国家公园之一,每年有900万——950万游人。位于田纳西州东部的大型游乐场多莱坞(Dollywood)的游客每年超过200万人。如果我们去大烟山和多莱坞旅游,就会发现几乎处处都是肥胖者。
1991年,美国没有任何一个州的肥胖人口超过20%。仅仅20多年间发生的变化无法用基因变化来解释。现在85%以上的美国人认为,肥胖是一种病。美国疾病控制预防中心(Centers for Disease Control and Prevention)和世界卫生组织的调查确认,肥胖是仅次于吸烟的第二大流行病,并将在10年内成为世界第一大流行病。(现代人的食物远远超过了实际需求。线粒体以氧气为原料,每天制造的ATP能量的重量占人体体重的一半,为人类制造能量的效率为20万倍)
现代饮食中,排名第一位的罪犯是糖。人类的基因因为无法处理过量的糖(碳水化合物),从而导致糖尿病。另一个重要罪犯是添加剂。2002年,埃里克·施洛瑟(Eric Schlosser,1959-)出版了大型调查报告《快餐民族:所有美国人食物的黑暗面》(Fast Food Nation: The Dark Side of the AllAmerican Meal)。这本书列举了很多数据,例如,麦当劳草莓奶昔由60多种添加剂构成,唯独没有任何草莓成分,含糖很多。又如,番茄酱(ketchup)的三分之一是糖。这本调查报告引起巨大轰动,美国涌现大量类似书籍,出现多部电影,批判反思现代饮食文化。
狩猎采集时代的田园牧歌,不可能时光倒流。(北美土著有一句古老格言:善等地球。它不是你父母给你的,它是你的孩子们借给你的。Treat the earth well. It was not given to you by your parents, but is loaned to you by your children.)
农业文化认为,向地球索取可以无穷无尽,尤其最近几个世纪的无限制扩张和掠夺几乎达到疯狂。可是,土地终有尽头,地球终有尽头。 农业文化发展进步的陈旧模式面临资源枯竭的致命挑战,继续维持已不可能。虽然我们无法回到农业以前的时代,但是狩猎采集时代的人类文化值得我们反思和借鉴。 人口与人,完全是两个不同的概念。托马斯·马尔萨斯(Thomas Robert Malthus,1766-1834)说:“人口的力量无限大于索取地球而求生存的人的力量。”
1900年,一场不文明的内战打响了,这是孟德尔的遗传学针对达尔文的自然选择的战争。大部分生物学家认为,这场战争的结果将是一个理论灭绝另一个理论。三个复活了孟德尔的科学家之一,胡戈·德弗里斯(Hugo de Vries)首先发明了突变理论(mutation theory),他认为物种起源是某些罕见的突变引起的。
摩尔根的小小的纽约实验室原本拥挤狭小得滑稽可笑,1928年,成为生物学的“重要人物”之后,他搬到了加利福尼亚州的洛杉矶的宽敞明亮的新实验室里,雄心勃勃地希望建立自己的理论体系,虽然他的果蝇实验和突变理论实际上只是追随别人的实验模式和理论。他在洛杉矶加州理工学院(California Institute of Technology)创建的生物系,先后培育出了7个诺贝尔奖得主。
1952年,更好的证据出现了。阿弗雷德·赫希(Alfred Day Hershey,1908-1997)和他的女助理玛莎·蔡斯(Martha Cowles Chase,1927-2003),在美国纽约的冷泉港实验室(Cold Spring Harbor Laboratory)利用病毒进行的著名的赫希——蔡斯实验(Hershey-Chase experiment),证实了DNA是遗传物质。
1953年,一年之后,剑桥大学的两个年轻人,詹姆斯·沃森和弗朗西斯·克里克终于搞清了DNA的奇特而稳定的化学分子结构。1954年,20个青年学者(代表20种氨基酸)组成RNA领带俱乐部(RNA Tie Club),讨论分析DNA→RNA→蛋白的遗传关系:DNA是双链,RNA是单链,DNA将遗传信息交给“信使”RNA,然后由RNA指令细胞制造蛋白。这个遗传信息的转达和表达过程,转瞬即逝,机理难以查明。DNA→RNA→蛋白的遗传制造过程中,当然也会出现错误,但是细胞通常会立刻修正这些错误,否则这些错误就永久留在DNA里遗传下去。
1958年,DNA结构的两个发现者之一,弗朗西斯·克里克(Francis Crick)发布了著名的“分子生物学中心法则”(Central dogma of molecular biology)。这个中心法则的主要含义是:DNA制造RNA制造蛋白质(DNA makes RNA makes protein)。
(1979年洛夫洛克出版了《盖亚:对地球生命的新看法》(Gaia: A New Look at Life on Earth),这是洛夫洛克出版的“盖亚理论”的第一本书。他出版了多部著作,如: Lovelock, James. Gaia: A New Look at Life on Earth. Oxford University Press, Oxford, England. 洛夫洛克的这个观念,其实并非全新的观念。)
加利福尼亚州的红杉(S e q u o i a gigantea)是生命的最好注解。这些巨树生长在树丛里,高度达到100米以上,寿命超过3 000年。红杉97%的组织是死的,主干和树皮已经死去,只有主干外表的细胞部分是活的。红杉的主干类似地球的岩石圈,只有岩石圈外表薄薄一层生物圈是活的。红杉的树皮类似大气层,保护着这层生物圈,并且进行生物学意义上非常重要的气体交换——二氧化碳和氧气的交换。 毫无疑问,红杉总体上是活的生命,我们不能只把红杉的外层称为红杉,其余部分视为死的木头。
2012年9月5日,人们又一次发现自己错了。 2012年9月5日开始,世界第一大媒体《时代》的一篇报道的题目本身就蕴含认错的含义:《垃圾基因:其实并非无用》(Junk DNA-Not So Useless After All)。这篇报道连续5天占据《时代》网络版头版位置。 这个消息,也是世界所有媒体的头版新闻。
YAP是Y染色体Alu多态性(Y Alu Polymorphism)的简称,Alu是Y染色体上长度约300碱基对(核苷酸)的一个区段,又称阿鲁元素(Alu element),这个无害的Alu重复地插入人类基因组的不同部位,插入模式已经超过100万种并遗传给后裔。约5万年前,一个男人体内的Y染色体上出现了这个300碱基对的区段并遗传给他的后裔。
奴隶制度的拥护者曾经认为,现代人分为很多物种和亚种,殖民者与奴隶不是一个物种。瑞典科学家卡尔·冯·林奈(Carl von Linne)最早提出这一体系。林奈是一个植物学家,他首先用拉丁文命名植物,随后扩展到动物。他把人类命名为智人(Homo sapiens)。他认为所有的人类属于同一物种——智人的不同亚种和地理种,他还认为人的种族是互不相同的、分别诞生的、多元发生的。这种思想起始于希腊时代的人类“多起源说”。
《物种起源》:原书全称《物种起源,通过自然选择的方式或在生存斗争保留优势种群的方式》(On the Origin of Species by Means of Natural Selection, or the Preservation of Favoured Races in the Struggle for Life)(注:这本书英文原名《种的起源》,日文译名也是《种的起源》,国内长期译为《物种起源》,本书沿袭旧译名)。
《人的由来》:原书全称叫作《人的由来,与性关系的选择》(The Descent of Man,and Selection in Relation to Sex)。
1960年代,人类学家的世界最高权威之一、美国体质人类学家协会(American Association of Physical Anthropologists)会长卡尔顿·库恩(Carleton Coon)发表了影响很大的两本著作:《种族的起源》(The Origin of Races)和《人的现存种族》(The Living Races of Man)。库恩在他的权威巨著中,把现代人类进一步细分为互不相同的五大亚种(实际是地理种): Australoid :澳大利亚人种(澳大利亚土著,又称棕种人); Caucasoid :高加索人种(欧洲——北非——西亚——中亚——南亚,又称白种人,虽然肤色不一); Negroid :尼格罗人种(非洲撒哈拉南部,东南亚小岛与山区,又称Congoid或黑种人); Capoid :开普敦人种(非洲南部,如布须曼人——桑人); Mongoloid :蒙古利亚人种(亚洲大部——北极圈——南北美洲——太平洋诸岛,又称黄种人)。
1987年1月,美国一个在读遗传学女博士丽贝卡·卡恩(Rebecca Cann)和她的同事们在英国《自然》(Nature)杂志发表了一篇论文:《线粒体DNA和人类的演化》(Mitochondrial DNA and Human Evolution)。论文认为:人类起源只有一个,这个起源可能在非洲,时间在20万年以内。尽管几乎不可思议,但DNA数据研究分析却证明了:今天所有的地球人都来自同一个共同祖先。
1987年,英国《自然》(Nature)杂志上发表的著名的mtDNA树:世界147个人的线粒体DNA计算分析。论文的三个署名作者与顺序:丽贝卡·卡恩 Rebecca Cann 马克·斯通尼金 Mark Stoneking 阿伦·威尔逊 Allan Wilson 这是一个里程碑
这就是“线粒体夏娃”的秘密。
1987年,丽贝卡·卡恩和她的同事发表这一结果之后,面对激烈的争议和质疑,他们又开始了一项新的研究。1987年9月,丽贝卡·卡恩和她的同事们又在英国《自然》(Nature)杂志发表了第二篇论文:《有争议的人类群体非洲起源》(DisputedAfrican origin of human populations),再次证实“线粒体夏娃”确实就在非洲。
但是,DNA的这些X射线的照片的形态太奇特了,他们两人很久都没有找到与之相吻合的DNA化学结构。最后,沃森和克里克采用了非常笨拙的“原始”方法,用一条一条的硬纸板和一片一片的金属片和金属丝,构建各式各样的模型,试图重现DNA的结构。他们最后发现有一种双螺旋模型完全吻合X射线的分布。这种模型很简单,像两个螺旋形的梯子扭结在一起。而且这种结构非常稳定——只有稳定的结构才能作为遗传物质。DNA正是一种极其稳定的化学结构,仅由4种核苷酸碱基构成的一种糖类骨架。这4种核苷酸的化学名称如下: A 腺嘌呤 C 胞嘧啶 G 鸟嘌呤 T 胸腺嘧啶
宗族母亲(宗族母亲必须有两个女儿,而不是一个女儿。母系祖先是这8个女人最晚近的共同祖先——她的母亲当然也是此后所有女人的母系祖先,但她母亲不是最晚近的,她本人才是。她的两个女儿也是后继的女人的母系祖先,但没有一个是所有这8个女人的共同母系祖先。也就是说,如果将上图视为一个宗族,只有标为MRCA的这个女人是宗族母亲。不论8个人还是800万人的宗族都适用这同一原则(MRCA,Most Recent Common Ancestor:最晚近的共同先祖))
1996年,美国的《历史频道》(The History Channel)在《最伟大的法老》节目中列举了15位埃及法老,其中第10位伟大法老是图坦卡蒙(Tutankhamun),他的在位时间大约是公元前1334-前1325年:八九岁即位,在位大约10年。这位图坦卡蒙法老,任何“伟大的功业”也没有干过。他的“伟大”之处仅仅在于他的陵墓在3 300多年里没有被盗,是唯一没有被盗的埃及法老陵墓。
帕博(Svante Paabo,1955-,曾经恢复埃及木乃伊的DNA和尼安德特人的DNA)1997年开始任马克斯·普朗克进化人类学研究所(Max Planck Institute for Evolutionary Anthropology)的遗传部主任。马普进化人类学研究所有五个部,属于德国马普研究院(Max Planck Society)。马普研究院有32位诺贝尔奖获得者,包括斯万特·帕博的父亲。1980年代,帕博和他的同行,包括找到“线粒体夏娃”的三个论文作者之一阿伦·威尔逊(Allan Wilson,1934-1991)等人,在德国和美国分别开始了古代DNA领域的研究。他们首先找出埃及木乃伊的DNA序列,然后很快转向化石。
1995年11月,在西班牙巴塞罗那召开的第二届欧洲群体历史会议(Second Euroconference on Population History)上,出现了一场激烈争辩。牛津大学的赛克斯在发言中用线粒体DNA的事实批驳了“欧洲人起源于中东农民”的主流观点。发言结束后的提问时间,“前进的浪潮”的支持者们提出各种意见,但是在DNA数据面前,却又无话可说。斯福扎也在会场,他没有多说什么。会议结束后的五年里,激烈的争论一直没有停止。这是又一场欧洲人的起源之战。
在科学界,巴塞罗那“欧洲群体历史会议”之类的国际会议可以宣布新的发现,但是会议的报告不是真正有效的,必须在科学期刊上发表。发表过程中,一批评审专家将对立题——结果——解释进行彻底的审查,称为同行评审。评审专家必须与作者没有任何利益冲突。牛津大学把报告送到《美国人类遗传学杂志》(American Journal of Human Genetics),受到非同寻常的严格审查。不仅要求对1995年发表的数学化的晦涩难懂的网络构建方法加上一个附录,作出进一步解释,还要求加上传统的群体比较表格。从巴塞罗那会议到论文发表,评审拖延长达8个月。当时世界各国的实验室和大学都在进行DNA的研究,没有统一标准,甚至互相保密,方法不同,DNA的表述和编号也不一致,这一切直到美国的人类基因组工程之后才逐步统一。
达尔文是一位性格平和、实事求是的博物学家,他喜欢观察,喜欢化石。他的名字前面的一长串各式各样的称号,都是后人添加的头衔。在《物种起源》中,达尔文甚至没用“进化”(evolution)一词,而是采用了“更改的后代”(descent with modification)一词,因为他认为,进化一词含有进步的含义,物种的遗传变化只是为了适应变化的环境,并没有进步或退步的含义。达尔文除了收集各个物种的标本,还收集了大量化石。但是,达尔文当时还无法分辨清楚这些化石,也没有条件进行统计学的分析。
但是,正如达尔文的推测,世界上化石最多的地方在非洲。1920年代,非洲的猿人化石开始大量出土,远远超过欧洲和亚洲。1921年,赞比亚发现第一个猿人化石。1922年,雷蒙德·达特(Raymond Dart,1893-1998)被任命为南非威特沃特斯兰德大学(University of the Witwatersrand)的人类学教授,开始组建一个人类学系。1924年,达特确认,在赞比亚发现的是迄今最古老的猿人化石。1959年,在距离赞比亚几千千米的肯尼亚,路易斯·李基(Louis Leakey)发现了一个175万年前的南方古猿(Australopithecus)。这一考古发现,将非洲地区的远古类人猿的生存年代延长了大约一倍。此后的考古发现,非洲人科生物的化石年代越来越久远,分布越来越广泛。
此后的几十年里,越来越多的非洲南方古猿(Southern Ape Man)的化石大量出土,其数量之大,超出世界上其他所有地方的总和。人类起源于非洲的理论,在事实面前逐渐被世界接受。
非洲南方古猿(Southern Ape Man)的年代逐渐向前延伸:300万年,400万年……最新发现的类似黑猩猩的猿人Ardipitbecus(地猿)进一步把非洲猿人的年代延伸到560万年前的中新世(Miocene)。但是,伯克利大学计算出来的“线粒体夏娃”这个现代智人的诞生时间,仅仅不到20万年,这到底是怎么一回事?(1974年11月,露西(Lucy)在埃塞俄比亚出土,她的年龄约20岁,生活年代约320万年前。露西属于南方古猿阿法种(Australopithecus afarensis),这种古猿与现代人的关系目前仍不清楚。露西被列为联合国世界文化遗产)
1994年,诺贝尔奖获得者沃特·吉尔伯特(Walter Gilbert)和罗布·多利特(Rob Dorit)、广濑明石(Hiroshi Akashi)在《科学》上发表了一篇奇特的论文。这篇论文的奇特在于:他们不是报道发现了什么,而是报道没有发现什么,论文的题目是《人类Y染色体在ZFY区段不存在多态性》(Absence ofpolymorphism at the ZFY locus on the human Y-chromosome)。这三个科学家希望能从世界不同地方采样的38个人的Y染色体上找到多态性,但是最终没有找到。他们感到非常惊讶,反复进行核实,结果还是找不到。也就是说,这38个人理论上来自同一个父亲。一位诺贝尔奖得主和两个生物专家,花费很大力气,发现了一个花天酒地的风流男人,他在全世界眠花宿柳,他生下的38个儿子又恰好被搜集到这场科学实验里来了。 这绝对不可能。
此前,列文庭自己也曾根据地理上的“血统”粗略划分了人类: 高加索人(Caucasians ,欧亚大陆的西部) 黑色非洲人(Black Africans,撒哈拉以南的非洲) 蒙古人(Mongoloids,亚洲的东部) 南亚土著(South Asian Aborigines,印度的南部) 美洲人(Amerinds,美洲) 大洋洲和澳大利亚土著(Oceanians and Australian Aborigines)
从效果上看,简约理论为我们提供了一个哲学的时间机器,使我们得以返回早已不复存在的时代,四处探索和欣赏。这个机理,令人陶醉。其实,达尔文也是这个理论的一个懵懵懂懂的早期附和者。赫胥黎(Huxley)曾经批评达尔文本人对于自己的信仰也是稀里糊涂的,他说:“natura non facit saltum(拉丁语:大自然不会产生飞跃)。”
斯福扎和爱德华兹检测了世界各地的15个人群血型频率,用电脑分析了频率检测结果——非洲人之间的差异最大,欧洲人和亚洲人之间频率比较集中。这是人类进化历史中的一个激动人心的清晰证据。斯福扎说,这个分析“只是找到一些感觉”(made some kind of sense)。此外,这个分析结果反映出基因频率的相似性:随着时间的推移,频率呈现有规律的变化。
1980年代,人类在细胞里发现了一个小的结构:线粒体(mitochondrion)。2000年代,人们终于知道,线粒体是十几亿年以前第一批复杂细胞进化过程中留存下来的一类细菌。也就是说,我们的单细胞先祖们曾经吞噬了一种古代细菌,因为这种细菌在细胞内部可以生产能量,最后,这种被吞噬的古代细菌从一种“寄生虫”演变成一座亚细胞能量工厂(sub-cellular power plant)。非常幸运的是,与细菌的基因组类似,线粒体基因组(mitochondrial genome或者mtDNA)只有一套复制品,也就是说,它们不会重组。光明和希望再次出现。
我们再来看看Y染色体。Y染色体活跃基因丧失的情况与线粒体类似,虽然平均每对染色体有1 500个活跃基因,但是Y染色体上只剩下21个活跃基因,其中一些基因还是重复的,随机复制的。更有趣的是,这21个活跃基因只参与一项工程——制造男性。其中一个基因决定性别,称为Y染色体性别决定区,缩写SRY(Sex-determining Region of the Y)。其他的活跃基因负责决定其他男性特征(例如男人的外貌、长相、行为举止等)。Y染色体上的其他基因什么功能也没有,被称为“垃圾DNA”(junk DNA)。这些“垃圾DNA”也许是生物学的垃圾,却是群体遗传学家的金砂。
1991年,斯坦福大学的斯福扎实验室里来了一个应聘的青年人,彼得·安德希尔(Peter Underhill)。安德希尔早年在特拉华大学(University of Delaware)从事海洋生物学研究并获得博士学位,后来到加利福尼亚州,转向研究酶在分子生物学中的应用。1980年代正是生物技术大发展的初期,硅谷是重组DNA的震中。如何用各种各样的酶切割基因——分离基因——黏合基因……各种生物技术与电脑技术相互辉映,电子和生物两大技术领域将旧金山湾区变成了一个朝气蓬勃的全球新兴技术中心。
其他灭绝的类人猿或直立人都没有能力渡过海洋。例如,爪哇直立人距离澳大利亚的连成一片的大陆只有大约100千米,但是它们从未到达澳大利亚。事实上,在澳大利亚从来没有发现任何灵长目的痕迹。只有现代智人具备渡过海洋的能力。当时属于旧石器时代,沿着海岸高速公路零零散散发现的旧石器时代的石器证明,当时非洲人的确是沿着海岸线走到澳大利亚的。虽然印度沿岸的证据沉睡在深深的海底,斯里兰卡的一个山洞(Fa Hien cave)里却出土了大量旧石器时代的石器,证实了远古沿海高速迁移的真实存在。澳大利亚的土著,正是已经部分沉入海洋之下的澳大拉西亚(Australasia)的最早的一批殖民者,在他们的文化里,至今保留着祖传的因素。直到今天,澳大利亚的土著还保留着用歌声呼唤非洲先祖的仪式。
如果群体的增长是停滞的或者缩减的,这个分布会呈现“来回拉锯”的形态,原因是遗传漂变(genetic drift)或自然选择导致某些血统的绝嗣和丧失。如果这种分布形态是平滑的,表示我们现代智人的人口增长速率很高。哈本丁和他的团队采集和分析了世界25个群体的线粒体DNA数据,发现呈现指数增长的群体多达23个。根据这一研究成果,他出版了一本著作《一万年的爆炸》(The 10000 Year Explosion)。亨利·哈本丁认为,这场人口大迁移起始于5万年前,这个时间与人类走出非洲的时间基本吻合。
D NA的复制过程如下: 由一批不同类型的小小复制机器——聚合酶(polymerases),先把双螺旋的两个链条打开,然后辛辛苦苦地分别复制两个链条的互补部分,分别形成另外两个DNA分子的双螺旋,使得一个DNA双螺旋变成了两个DNA双螺旋。这里只有一个简单的不可侵犯的法则:A永远配对T;C永远配对G。(1984年,遗传学家阿莱克·杰弗里斯(Alec Jeffreys,1950-)发现:3-30个碱基对的短核苷酸序列,在基因组里可以重复20-100次。他把这种重复序列组称为“微卫星”,或随机重复变量(VNTRs,variable number of tandem repeats)。人类基因组中这些区段的数量和位置,每个人都不一样。在人类的旅程的探索中,这种“微卫星”技术大量运用,找出了各地的人类群体差别以及群体之内的个体之间的微小差异)
1787年,美国总统托马斯·杰弗逊(Thomas Jefferson)在他的《弗吉尼亚州笔记》(Notes on the State of Virginia)里写道: ……虽然亚洲与美洲是完全分离的,但是,中间只有一个狭窄的海峡……美洲印第安人与亚洲东部的居民之间相似的外貌使我们产生一个猜测,要么前者是后者的后裔,要么后者是前者的后裔……
事情并未到此结束。1986年:著名的《自然》(Nature)刊登了巴西考古学家尼埃德·古伊登(Niede Guidon,1933-)的一篇令人震惊的文章:《碳14显示人类3.2万年前在美洲》(Carbon-14 dates point to man in the Americas 32 000 years ago)。这篇文章介绍了在巴西东北部皮奥伊州(Piaui)的大批洞穴发现的史前遗迹和各种壁画。这些壁画总数超过三万处,除了远古时代的礼仪、舞蹈、狩猎以外,还有最后一次冰河期以前灭绝的动物雕齿兽(Glyptodon)、巨型犰狳(Armadillo)等动物。这里出土了大量陶器,还有绘制的世界最早的船只。这篇文章在美洲迁移史的研究中引起了轩然大波。这些历史遗迹的具体时间,至今仍然在争议中。
1999年,Fabricio Santos和Chris Tyler-Smith在牛津大学,Tanya Karafet和 Mike Hammer在亚利桑那大学(University of Arizona),分别独立地报告,M3的祖先是Y染色体上的一个未加定义的核苷酸改变,这个基因标记叫作92R7。他们发现从欧洲到印度的整个欧亚大陆都有92R7。这个92R7外加其他核苷酸变化,共同证实西伯利亚是美洲土著的来源。这一结论也佐证了线粒体DNA研究的结果。但是,研究者却难以确定92R7血统的年龄,因为这个基因标记太普遍了。
1786 年,在英属印度担任法官的语言学家威廉姆·琼斯爵士(Sir William Jones,1746-1794,发现梵语与拉丁语和希腊语非常相似),在大量研究分析的基础上提出印欧语系(Indo-European language family)的概念,即欧洲到印度的大片地区所有的语言有一个共同起源。这种假设,最终得到广泛承认。
波利尼西亚是太平洋中央几千个岛屿的统称,波利尼西亚人是这些岛屿原住民的统称,从夏威夷的土著到新西兰的毛利人(Maori people)都属于这一范畴。波利尼西亚(Polynesia)一词源自希腊语:poly意为众多,nesoi意为岛屿。1756年法国作家Charles de Brosses(1706-1777)第一次使用这个词,当时泛指太平洋上的所有岛屿。现在的波利尼西亚(Polynesia)的范围也没有严格界定:从美国洛杉矶出发,飞向新西兰首都奥克兰经过的海域就是波利尼西亚,距离约1.2万千米,飞行约14个小时。飞机下面一望无际的辽阔海域散布着数以千计的岛屿。
因为收集的人类DNA样本越多越好,于是,这个计划向尽可能多的人群开放,包括那些背景复杂、遗传形态极其难以识别的群体。任何愿意了解自己DNA的人都可以购买一套自我检测套件——基因图谱工程公众参与套件(Genographic Project Public Participation Kit),把受检者引人入胜的DNA故事加入工程设立在世界各地的11个收集检测中心,最后,所有数据通过互联网进入数据库汇总分析计算。
欧洲的王室家族的近亲婚姻,最典型的例子是哈布斯堡王朝(House of Habsburg)。这个王朝是欧洲历史上最有权势的王朝,起源于奥地利、匈牙利,他们通过婚姻关系扩大政治联盟。哈布斯堡王朝的鼎盛时代几乎联姻到了欧洲的每一个王室,使得16-18世纪的欧洲王室之间的血缘关系极其接近,遗传学效果使他们几乎成为一个“小村庄”。这个王朝不仅一代又一代地遗传财富和权势,也遗传基因标记和各种生理缺陷。按照遗传学意义,这个王朝最后变为同系交配或同族交配(endogamous),越来越差的王室后代正是哈布斯堡王朝最后土崩瓦解的重要原因之一。
M168又被称为“欧亚大陆亚当”(Eurasian Adam)或“走出非洲亚当”(Out of Africe Adam),这个男人的Y染色体突变发生的时间为6万——7.9万年前,地点在东非的埃塞俄比亚——苏丹一带。我们不知道M168是什么人,有些学者认为他“可能是一夫多妻制度下的一位酋长”。走出非洲的男性并非只有M168的后裔,但是M168是迄今为止唯一没有断绝的男性Y染色体血统。
基因图谱工程采用的单倍群编码。图片来源:美国国家地理协会
M 168重要的后裔也有3个人,即M130、YAP和M89。约6万年前,第一批走出非洲的人类是M168的后裔M130。一部分M130沿着海岸线一直走到澳大利亚,还有一些M130留在印度次大陆——东南亚地区,他们继续北上进入亚洲的东部,即青藏高原——中国内地——蒙古——韩国——日本等地,还有一些人进入了北美洲。
第一个DNA证据,正是来自这位大学里的图书管理员韦鲁曼迪。他的一个基因标记名叫RPS4Y,这个缩写名称的全称是Ribsomal Protein S4 on the Y chromosome(Y染色体上的核糖体蛋白质S4)。这个RPS4Y,现在简称M130:Y染色体上发现的第130个基因标记。在印度南部的人群中,M130的频率仅约为5%,包括卡拉尔群体(Kallar)。但是在澳大利亚土著中,M130却成为主导标记,超过50%。在东南亚约为20%。在印度的北部地区也发现了M130。
2010年2月5日,《新西兰先驱》(The New Zealand Herald)刊登了一篇题为《达尔文家族DNA的非洲起源》(Darwin family DNA shows African origin)的报道。1986年,达尔文的直系后裔克里斯·达尔文(Chris Darwin)移居澳大利亚,居住在悉尼西边的Blue Mountains。2010年,克里斯·达尔文接受人类基因图谱工程的DNA分析,证实达尔文的家族约四万年前走出非洲,路线为中东——中亚——欧洲,最后一次冰河时代辗转进入西班牙,然后北上迁移到英国。