Press "Enter" to skip to content

“主要的网站正在阻止AI爬虫访问他们的内容”

.fav_bar { float:left; border:1px solid #a7b1b5; margin-top:10px; margin-bottom:20px; } .fav_bar span.fav_bar-label { text-align:center; padding:8px 0px 0px 0px; float:left; margin-left:-1px; border-right:1px dotted #a7b1b5; border-left:1px solid #a7b1b5; display:block; width:69px; height:24px; color:#6e7476; font-weight:bold; font-size:12px; text-transform:uppercase; font-family:Arial, Helvetica, sans-serif; } .fav_bar a, #plus-one { float:left; border-right:1px dotted #a7b1b5; display:block; width:36px; height:32px; text-indent:-9999px; } .fav_bar a.fav_print { background:url(‘/images/icons/print.gif’) no-repeat 0px 0px #FFF; } .fav_bar a.fav_print:hover { background:url(‘/images/icons/print.gif’) no-repeat 0px 0px #e6e9ea; } .fav_bar a.mobile-apps { background:url(‘/images/icons/generic.gif’) no-repeat 13px 7px #FFF; background-size: 10px; } .fav_bar a.mobile-apps:hover { background:url(‘/images/icons/generic.gif’) no-repeat 13px 7px #e6e9ea; background-size: 10px} .fav_bar a.fav_de { background: url(/images/icons/de.gif) no-repeat 0 0 #fff } .fav_bar a.fav_de:hover { background: url(/images/icons/de.gif) no-repeat 0 0 #e6e9ea } .fav_bar a.fav_acm_digital { background:url(‘/images/icons/acm_digital_library.gif’) no-repeat 0px 0px #FFF; } .fav_bar a.fav_acm_digital:hover { background:url(‘/images/icons/acm_digital_library.gif’) no-repeat 0px 0px #e6e9ea; } .fav_bar a.fav_pdf { background:url(‘/images/icons/pdf.gif’) no-repeat 0px 0px #FFF; } .fav_bar a.fav_pdf:hover { background:url(‘/images/icons/pdf.gif’) no-repeat 0px 0px #e6e9ea; } .fav_bar a.fav_more .at-icon-wrapper{ height: 33px !important ; width: 35px !important; padding: 0 !important; border-right: none !important; } .a2a_kit { line-height: 24px !important; width: unset !important; height: unset !important; padding: 0 !important; border-right: unset !important; border-left: unset !important; } .fav_bar .a2a_kit a .a2a_svg { margin-left: 7px; margin-top: 4px; padding: unset !important; }

任何您可以通过网络浏览器访问的页面也可以被爬虫“抓取”,它的工作原理与浏览器相似,但是将内容存储在数据库中而不是显示给用户。 ¶ 来源:Annelise Capossela/Axios

根据AI内容检测器Originality.AI的最新数据,全球前1000个网站中有近20%的网站正在阻止用于AI服务的爬虫程序收集网络数据。

为什么重要:在没有明确法律或监管规定规范AI使用受版权保护的材料的情况下,大小网站都在采取行动。

新闻动态:OpenAI于8月初推出了其GPTBot爬虫,并声明收集到的数据“可能用于改进未来的模型”,承诺将排除付费内容,并指导网站如何阻止该爬虫。

此后不久,包括纽约时报、路透社和CNN在内的几家知名新闻网站开始阻止GPTBot,此后还有更多网站效仿(Axios也在其中)。

来自 Axios 查看完整文章

Leave a Reply

Your email address will not be published. Required fields are marked *