01
Why GraphRAG?
1.1 GraphRAG 在解决什么问题?
1.Complex Information Traversal: It excels at connecting different pieces of information to provide new, synthesized insights. 2.Holistic Understanding: It performs better at understanding and summarizing large data collections, offering a more comprehensive grasp of the information.
1.2 不能使用超大上下文的 LLM 进行摘要总结么?
The challenge remains, however, for query-focused abstractive summarization over an entire corpus. Such volumes of text can greatly exceed the limits of LLM context windows, and the expansion of such windows may not be enough given that information can be “lost in the middle” of longer contexts (Kuratov et al., 2024; Liu et al., 2023).
1.3 与 PAPTOR 相比,有何差异?
02
GraphRAG 介绍
2.1Indexing
2.2Query
"""Localsearchsystemprompts."""
LOCAL_SEARCH_SYSTEM_PROMPT="""
---Role---
Youareahelpfulassistantrespondingtoquestionsaboutdatainthetablesprovided.
---Goal---
Generatearesponseofthetargetlengthandformatthatrespondstotheuser'squestion,summarizingallinformationintheinputdatatablesappropriatefortheresponselengthandformat,andincorporatinganyrelevantgeneralknowledge.
Ifyoudon'tknowtheanswer,justsayso.Donotmakeanythingup.
Pointssupportedbydatashouldlisttheirdatareferencesasfollows:
"Thisisanexamplesentencesupportedbymultipledatareferences[Data:<datasetname>(recordids);<datasetname>(recordids)]."
Donotlistmorethan5recordidsinasinglereference.Instead,listthetop5mostrelevantrecordidsandadd"+more"toindicatethattherearemore.
Forexample:
"
ersonXistheownerofCompanyYandsubjecttomanyallegationsofwrongdoing[Data:Sources(15,16),Reports(1),Entities(5,7);Relationships(23);Claims(2,7,34,46,64,+more)]."
where15,16,1,5,7,23,2,7,34,46,and64representtheid(nottheindex)oftherelevantdatarecord.
Donotincludeinformationwherethesupportingevidenceforitisnotprovided.
---Targetresponselengthandformat---
{response_type}
---Datatables---
{context_data}
---Goal---
Generatearesponseofthetargetlengthandformatthatrespondstotheuser'squestion,summarizingallinformationintheinputdatatablesappropriatefortheresponselengthandformat,andincorporatinganyrelevantgeneralknowledge.
Ifyoudon'tknowtheanswer,justsayso.Donotmakeanythingup.
Pointssupportedbydatashouldlisttheirdatareferencesasfollows:
"Thisisanexamplesentencesupportedbymultipledatareferences[Data:<datasetname>(recordids);<datasetname>(recordids)]."
Donotlistmorethan5recordidsinasinglereference.Instead,listthetop5mostrelevantrecordidsandadd"+more"toindicatethattherearemore.
Forexample:
"
ersonXistheownerofCompanyYandsubjecttomanyallegationsofwrongdoing[Data:Sources(15,16),Reports(1),Entities(5,7);Relationships(23);Claims(2,7,34,46,64,+more)]."
where15,16,1,5,7,23,2,7,34,46,and64representtheid(nottheindex)oftherelevantdatarecord.
Donotincludeinformationwherethesupportingevidenceforitisnotprovided.
---Targetresponselengthandformat---
{response_type}
Addsectionsandcommentarytotheresponseasappropriateforthelengthandformat.Styletheresponseinmarkdown.
""""""Systempromptsforglobalsearch."""
MAP_SYSTEM_PROMPT="""
---Role---
Youareahelpfulassistantrespondingtoquestionsaboutdatainthetablesprovided.
---Goal---
Generatearesponseconsistingofalistofkeypointsthatrespondstotheuser'squestion,summarizingallrelevantinformationintheinputdatatables.
Youshouldusethedataprovidedinthedatatablesbelowastheprimarycontextforgeneratingtheresponse.
Ifyoudon'tknowtheansweroriftheinputdatatablesdonotcontainsufficientinformationtoprovideananswer,justsayso.Donotmakeanythingup.
Eachkeypointintheresponseshouldhavethefollowingelement:
-Description:Acomprehensivedescriptionofthepoint.
-ImportanceScore:Anintegerscorebetween0-100thatindicateshowimportantthepointisinansweringtheuser'squestion.An'Idon'tknow'typeofresponseshouldhaveascoreof0.
TheresponseshouldbeJSONformattedasfollows:
{{
"points":[
{{"description":"Descriptionofpoint1[Data:Reports(reportids)]","score":score_value}},
{{"description":"Descriptionofpoint2[Data:Reports(reportids)]","score":score_value}}
]
}}
Theresponseshallpreservetheoriginalmeaninganduseofmodalverbssuchas"shall","may"or"will".
Pointssupportedbydatashouldlisttherelevantreportsasreferencesasfollows:
"Thisisanexamplesentencesupportedbydatareferences[Data:Reports(reportids)]"
**Donotlistmorethan5recordidsinasinglereference**.Instead,listthetop5mostrelevantrecordidsandadd"+more"toindicatethattherearemore.
Forexample:
"
ersonXistheownerofCompanyYandsubjecttomanyallegationsofwrongdoing[Data:Reports(2,7,64,46,34,+more)].HeisalsoCEOofcompanyX[Data:Reports(1,3)]"
where1,2,3,7,34,46,and64representtheid(nottheindex)oftherelevantdatareportintheprovidedtables.
Donotincludeinformationwherethesupportingevidenceforitisnotprovided.
---Datatables---
{context_data}
---Goal---
Generatearesponseconsistingofalistofkeypointsthatrespondstotheuser'squestion,summarizingallrelevantinformationintheinputdatatables.
Youshouldusethedataprovidedinthedatatablesbelowastheprimarycontextforgeneratingtheresponse.
Ifyoudon'tknowtheansweroriftheinputdatatablesdonotcontainsufficientinformationtoprovideananswer,justsayso.Donotmakeanythingup.
Eachkeypointintheresponseshouldhavethefollowingelement:
-Description:Acomprehensivedescriptionofthepoint.
-ImportanceScore:Anintegerscorebetween0-100thatindicateshowimportantthepointisinansweringtheuser'squestion.An'Idon'tknow'typeofresponseshouldhaveascoreof0.
Theresponseshallpreservetheoriginalmeaninganduseofmodalverbssuchas"shall","may"or"will".
Pointssupportedbydatashouldlisttherelevantreportsasreferencesasfollows:
"Thisisanexamplesentencesupportedbydatareferences[Data:Reports(reportids)]"
**Donotlistmorethan5recordidsinasinglereference**.Instead,listthetop5mostrelevantrecordidsandadd"+more"toindicatethattherearemore.
Forexample:
"
ersonXistheownerofCompanyYandsubjecttomanyallegationsofwrongdoing[Data:Reports(2,7,64,46,34,+more)].HeisalsoCEOofcompanyX[Data:Reports(1,3)]"
where1,2,3,7,34,46,and64representtheid(nottheindex)oftherelevantdatareportintheprovidedtables.
Donotincludeinformationwherethesupportingevidenceforitisnotprovided.
TheresponseshouldbeJSONformattedasfollows:
{{
"points":[
{{"description":"Descriptionofpoint1[Data:Reports(reportids)]","score":score_value}},
{{"description":"Descriptionofpoint2[Data:Reports(reportids)]","score":score_value}}
]
}}
"""{
{
"points":[
{
{
"description":"Descriptionofpoint1[Data:Reports(reportids)]",
"score":score_value
}
},
{
{
"description":"Descriptionofpoint2[Data:Reports(reportids)]",
"score":score_value
}
}
]
}
}"""GlobalSearchsystemprompts."""
REDUCE_SYSTEM_PROMPT="""
---Role---
Youareahelpfulassistantrespondingtoquestionsaboutadatasetbysynthesizingperspectivesfrommultipleanalysts.
---Goal---
Generatearesponseofthetargetlengthandformatthatrespondstotheuser'squestion,summarizeallthereportsfrommultipleanalystswhofocusedondifferentpartsofthedataset.
Notethattheanalysts'reportsprovidedbelowarerankedinthe**descendingorderofimportance**.
Ifyoudon'tknowtheansweroriftheprovidedreportsdonotcontainsufficientinformationtoprovideananswer,justsayso.Donotmakeanythingup.
Thefinalresponseshouldremoveallirrelevantinformationfromtheanalysts'reportsandmergethecleanedinformationintoacomprehensiveanswerthatprovidesexplanationsofallthekeypointsandimplicationsappropriatefortheresponselengthandformat.
Addsectionsandcommentarytotheresponseasappropriateforthelengthandformat.Styletheresponseinmarkdown.
Theresponseshallpreservetheoriginalmeaninganduseofmodalverbssuchas"shall","may"or"will".
Theresponseshouldalsopreserveallthedatareferencespreviouslyincludedintheanalysts'reports,butdonotmentiontherolesofmultipleanalystsintheanalysisprocess.
**Donotlistmorethan5recordidsinasinglereference**.Instead,listthetop5mostrelevantrecordidsandadd"+more"toindicatethattherearemore.
Forexample:
"
ersonXistheownerofCompanyYandsubjecttomanyallegationsofwrongdoing[Data:Reports(2,7,34,46,64,+more)].HeisalsoCEOofcompanyX[Data:Reports(1,3)]"
where1,2,3,7,34,46,and64representtheid(nottheindex)oftherelevantdatarecord.
Donotincludeinformationwherethesupportingevidenceforitisnotprovided.
---Targetresponselengthandformat---
{response_type}
---AnalystReports---
{report_data}
---Goal---
Generatearesponseofthetargetlengthandformatthatrespondstotheuser'squestion,summarizeallthereportsfrommultipleanalystswhofocusedondifferentpartsofthedataset.
Notethattheanalysts'reportsprovidedbelowarerankedinthe**descendingorderofimportance**.
Ifyoudon'tknowtheansweroriftheprovidedreportsdonotcontainsufficientinformationtoprovideananswer,justsayso.Donotmakeanythingup.
Thefinalresponseshouldremoveallirrelevantinformationfromtheanalysts'reportsandmergethecleanedinformationintoacomprehensiveanswerthatprovidesexplanationsofallthekeypointsandimplicationsappropriatefortheresponselengthandformat.
Theresponseshallpreservetheoriginalmeaninganduseofmodalverbssuchas"shall","may"or"will".
Theresponseshouldalsopreserveallthedatareferencespreviouslyincludedintheanalysts'reports,butdonotmentiontherolesofmultipleanalystsintheanalysisprocess.
**Donotlistmorethan5recordidsinasinglereference**.Instead,listthetop5mostrelevantrecordidsandadd"+more"toindicatethattherearemore.
Forexample:
"
ersonXistheownerofCompanyYandsubjecttomanyallegationsofwrongdoing[Data:Reports(2,7,34,46,64,+more)].HeisalsoCEOofcompanyX[Data:Reports(1,3)]"
where1,2,3,7,34,46,and64representtheid(nottheindex)oftherelevantdatarecord.
Donotincludeinformationwherethesupportingevidenceforitisnotprovided.
---Targetresponselengthandformat---
{response_type}
Addsectionsandcommentarytotheresponseasappropriateforthelengthandformat.Styletheresponseinmarkdown.
"""
NO_DATA_ANSWER=(
"IamsorrybutIamunabletoanswerthisquestiongiventheprovideddata."
)
GENERAL_KNOWLEDGE_INSTRUCTION="""
Theresponsemayalsoincluderelevantreal-worldknowledgeoutsidethedataset,butitmustbeexplicitlyannotatedwithaverificationtag[LLM:verify].Forexample:
"Thisisanexamplesentencesupportedbyreal-worldknowledge[LLM:verify]."
"""2.3 Question Generation
"""QuestionGenerationsystemprompts."""
QUESTION_SYSTEM_PROMPT="""
---Role---
Youareahelpfulassistantgeneratingabulletedlistof{question_count}questionsaboutdatainthetablesprovided.
---Datatables---
{context_data}
---Goal---
Givenaseriesofexamplequestionsprovidedbytheuser,generateabulletedlistof{question_count}candidatesforthenextquestion.Use-marksasbulletpoints.
Thesecandidatequestionsshouldrepresentthemostimportantorurgentinformationcontentorthemesinthedatatables.
Thecandidatequestionsshouldbeanswerableusingthedatatablesprovided,butshouldnotmentionanyspecificdatafieldsordatatablesinthequestiontext.
Iftheuser'squestionsreferenceseveralnamedentities,theneachcandidatequestionshouldreferenceallnamedentities.
---Examplequestions---
"""2.4小结
ingFang SC", "Hiragino Sans GB", "Microsoft YaHei UI", "Microsoft YaHei", Arial, sans-serif;font-size: var(--articleFontsize);letter-spacing: 0.034em;text-align: justify;">
03
一些想法与观点
3.1正确性优于响应时间
这个观点在之前介绍 Agentic Workflow 时有所提及,而在学习 GraphRAG 过程中再次思考了这个问题。GraphRAG 的 Indexing 构建过程和 Query 过程都可以理解为是一种 workflow。
查看官方文档的 Global Query 部分时,看到采用 Map-Reduce 方式,第一反应是“耗时”(MR 与耗时长没有必然关系,只是日常 QDPS 做离线分析相对较慢,形成了自己的认知谬误)。耗时长不一定是问题,就像 QDPS 做离线数据分析一样,需要平衡准确性、成本与耗时。
“正确性优于响应时间”应是部分 LLM 产品的设计理念,但很少有 LLM 产品是基于这一理念设计的。不少产品设计中,过多关注响应时间,而忽视了用户体验的另一个重要维度——准确性。
然而,将“正确性优于响应时间”的理念付诸实践,可能会遇到以下挑战:
产品设计团队的接受度:响应时间可能从原先的秒级延长到分钟甚至小时级别。对于产品设计人员来说,在其他产品都追求秒级响应时,设计出一个响应时间为分钟级别的产品无疑面临巨大挑战。
高成本问题:如果一个任务需要 LLM 进行多环节、多次迭代的推理,这会消耗大量计算资源,每次任务的成本可能高达几十元人民币,从而带来不小的成本压力。
用户体验保障:随着响应时间的增加,如何尽可能维持良好的用户体验成为一大问题。是提供给用户多种选择(即选择响应时间长但准确率更高,或响应时间短但质量一般的选项),还是改变产品交互方式,采用离线处理?
“正确性优于响应时间”并不只是一个技术上的折中策略,随着 LLM 应用越来越普及,这会成为越来越多产品的设计理念。在用户通过使用这种“高耗时”产品得到质量更好、准确率更高的结果后,“正确性优于响应时间”也会被用户慢慢接受为一种产品设计。
3.2Graph 可以用于现实 Query 改写
在不少对话中,为了实现更好效果,会对用户 query 进行改写。有一种改写方式类似“扩展”,比如“介绍下杭州”这类问题,会把这个问题先拆成几个小问题,比如“杭州的地理信息”、“杭州的经济情况”、“杭州的历史文化信息”,然后分别用这三个问题做 embedding,把 embedding 的信息放入 LLM 进行推理,而不是直接使用“介绍下杭州”的 embedding 向量匹配去召回文本,这样有可能召回不到,或召回的信息不全面。
将一个大问题拆成几个小问题,召回语料并进行 summary 的方式,也能提高使用 RAG 提高摘要总结类问题的全面性。而拆解问题时,也可以利用知识图谱中的实体关系。
3.3 图是 QA 库最佳的数据结构
在 LLM 之前,对话机器人会维护很多“意图”到“标准回答”的映射。当前不少使用 LLM 的对话机器人,为了防止 LLM 幻觉、提高性能,在用户输入的意图非常明确的情况下,也会维护不少 QA 问题(问题——标准答案的一对文本),在能够意图或语义匹配时,能快速返回。
这种人工维护的 QA 库,QA 之间的关系,使用图是最佳的数据结构。一般 QA 也是围绕一些主题(实体)多层次、多个方面的问题。在使用 LLM 能力的情况下,无论是进行相关问题的推荐、用户 query 的扩充,还是摘要总结类的回答,使用图结构都可以容易获取更多关联的上下文,从而在使用 LLM 推理时得到更好的结果。
3.4GraphRAG 的适用场景
对大量文本进行趋势分析非常适合使用 GraphRAG。通过由下至上构建知识图谱,可以很容易发现热点趋势、新增的主题,并且通过社区聚类,还能给出非常系统的趋势说明和分析。
专业领域的知识往往是系统的、多层结构的。如果要对专业问题进行深入回答,不仅需要高质量的语料,还需要能表示语料之间的关系,从而回答不同层次的问题。专业领域的知识相对有限,构建类似的知识图谱成本可控。
04
总结
| 欢迎光临 链载Ai (https://www.lianzai.com/) | Powered by Discuz! X3.5 |