Lemmy.VG
  • Communities
  • Create Post
  • heart
    Support Lemmy
  • search
    Search
  • Login
  • Sign Up
irelephant [he/him]🍭@lemm.ee to TechTakes@awful.systemsEnglish · 1 day ago

Ai scraping is an effective DDoS on the entire interent

pod.geraspora.de

external-link
message-square
20
fedilink
84
external-link

Ai scraping is an effective DDoS on the entire interent

pod.geraspora.de

irelephant [he/him]🍭@lemm.ee to TechTakes@awful.systemsEnglish · 1 day ago
message-square
20
fedilink
Excerpt from a message I just posted in a #diaspora team internal f...
pod.geraspora.de
external-link
Excerpt from a message I just posted in a #diaspora team internal forum category. The context here is that I recently get pinged by slowness/load spikes on the diaspora* project web infrastructure (Discourse, Wiki, the project website, ...), and looking at the traffic logs makes me impressively angry. In the last 60 days, the diaspora* web assets received 11.3 million requests. That equals to 2.19 req/s - which honestly isn't that much. I mean, it's more than your average personal blog, but nothing that my infrastructure shouldn't be able to handle. However, here's what's grinding my fucking gears. Looking at the top user agent statistics, there are the leaders: 2.78 million requests - or 24.6% of all traffic - is coming from Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot). 1.69 million reuqests - 14.9% - Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (Amazonb...
  • db0@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    15
    ·
    23 hours ago

    It’s a constant cat and mouse atm. Every week or so, we get another flood of scraping bots, which force us to triangulate which fucking DC IP range we need to start blocking now. If they ever start using residential proxies, we’re fucked.

    • irelephant [he/him]🍭@lemm.eeOP
      link
      fedilink
      English
      arrow-up
      9
      ·
      23 hours ago

      I have a tiny neocities website which gets thousands of views a day, there is no way that anyone is viewing it often enough for that to be organic.

      • db0@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        11
        ·
        23 hours ago

        quickly, add some ad revenue :P

        • 𝕸𝖔𝖘𝖘@infosec.pub
          link
          fedilink
          English
          arrow-up
          6
          ·
          18 hours ago

          From ai vendors. Let them pay you for scraping you lol

    • self@awful.systems
      link
      fedilink
      English
      arrow-up
      8
      ·
      22 hours ago

      at least OpenAI and probably others do currently use commercial residential proxying services, though reputedly only if you make it obvious you’re blocking their scrapers, presumably as an attempt on their end to limit operating costs

      • MonkderVierte@lemmy.ml
        link
        fedilink
        English
        arrow-up
        2
        ·
        8 hours ago

        They have a botnet on residential devices?

        • froztbyte@awful.systems
          link
          fedilink
          English
          arrow-up
          3
          ·
          8 hours ago

          the term of art is “residential proxy” and there’s a ton of them

          for example: it’s the flipside of Bright’s free VPN service - through Bright Data they sell people access proxied via some user’s connection

          • irelephant [he/him]🍭@lemm.eeOP
            link
            fedilink
            English
            arrow-up
            1
            ·
            3 hours ago

            And companies like honey that pay you (a pittance) to proxy people’s requests to porn sites.

      • db0@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        5
        ·
        22 hours ago

        Oh never heard of that. I have blocked their scrapers via agents but I haven’t felt residential proxy pain.

        • Opus DEI@mastodon.me.uk
          link
          fedilink
          arrow-up
          14
          ·
          22 hours ago

          @db0 @self Residential Proxy Pain are playing at the Dublin Castle in Camden this Friday, £4 advance, £5 on the door

        • self@awful.systems
          link
          fedilink
          English
          arrow-up
          8
          ·
          22 hours ago

          here’s a mastodon post and linked blog post with some details on what currently sets it off

          • db0@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            3
            ·
            10 hours ago

            PS: Looks like that sync issue between our instances is resolved now?

          • db0@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            6
            ·
            19 hours ago

            Daym, I should set me up some iocane as well I think

TechTakes@awful.systems

techtakes@awful.systems

Subscribe from Remote Instance

Create a post
You are not logged in. However you can subscribe from another Fediverse account, for example Lemmy or Mastodon. To do this, paste the following into the search field of your instance: [email protected]

Big brain tech dude got yet another clueless take over at HackerNews etc? Here’s the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

Visibility: Public
globe

This community can be federated to other instances and be posted/commented in by their users.

  • 371 users / day
  • 1.23K users / week
  • 2.33K users / month
  • 5.33K users / 6 months
  • 1 local subscriber
  • 1.86K subscribers
  • 717 Posts
  • 20K Comments
  • Modlog
  • mods:
  • David Gerard@awful.systems
  • BE: 0.19.5
  • Modlog
  • Legal
  • Instances
  • Docs
  • Code
  • join-lemmy.org