No, your paranoia-outage is the true "national security threat" - a threat to freedom. Stop fanning the flames and calling for more government intervention.
I'm a bit on the fence about this, but leaning towards the "bad luck". I'm sure there's a large swathe of nuance that I'm missing, but my simplistic view is: If they don't want to be scraped then they don't get included in "the thing" which, at minimum, is a data point for consumers to make consumer decisions about.
It strongly depends on how scraped data is being used.
If your 'cat' hasn't been able to catch their 'mouse' then the cat needs to get smarter, or look for alternative sources of mice, or the cat should be considered 'unviable'.
Have you approached them to get access to their data? Have you explained to them how your service can benefit their business?
(I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect, but then I'm also an 'information wants to be free' kinda person, but the Internet is increasingly an untrustworthy place, so security is overruling narrative).
That's cutting the tether from a _lot_ (there's no way to overstate this) of useful information, but that's potentially one of the great things about the open AI/LLM models, is that all that info is baked in there.
Disclaimer: As far as I understand it. Please educate me if I'm way off the mark.
Also, how much of the content of reddit is in Common Crawl? (same disclaimer applies to this comment)
I also understand 'the archival mindset', I hoard a bunch of data. But I also understand the logarithmic graph of futility.
Government website data should be openly available one way or another (at least within the country, and with reasonable security provision against obvious maliciousness).
If it has to be scraped then there may be other problems (which include lack of resources to make the data API accessible).
Anything connected to network can become a national security threat. IoT devices, internet connected smart appliances and all. Computers that run some OS is least of the worry because of the various layers of protection that can be added to them. But it is not the case with the IoT devices or SMART(!) devices where, if the firmware or any app ecosystem gets hacked / cracked / modified or sideloaded without any user intervention, could turn into a large zombie C2C botnet that can cause havoc.
Proxying is least of the worries!
When manufacturers bring out the smart features that require constant data collection, monitoring (obviously in the name of service quality, troubleshooting, firmware / app updates and subscription of services (!!)), it is oblivious to the fact that security and privacy at a personal / state / national level is never considered (for various reasons such as the design, profit interests and overheads / compliance perspectives and what not!) to be part of the ecosystem. Bad actors with sufficient knowledge and tools could exploit and profit from it.
It has become more of a planned obsolescence / profit / greed based industrial / corporate culture in rolling out unwanted stuff / monitoring and controlling as a feature that gets exploited to the core.
Honesntly, I don't have an answer for that, but have some thoughts that I wanted to share.
I do not want a cleaning robot or a water purifier or a printer or a TV or a refrigerator or an oven or a smart light bulb or a smart lock or even my car to have undisclosed, unwarranted connectivity to internet for whatever reasons.
I do not care if it is the manufacture or the service provider or whomsoever it may be!
I want all of the network connectivity options to be explained with all the controlling options and boundaries and data collection that happens on any of the connected medium to be as much transparent with the option of preventing or restricting what one does not want to go out of the device.
Will anyone do that?
It ill never happen! Sadly!
This problem will balloon further without adequate controls and killing the networking at a ground level would only be the only solution!
But, in the case of a connected device having an inbuilt connectivity option (like the m2m based options) in a smart vehicle, even that is not possible!
Ad networks like meta ads, google ads use residential IPs + small LLM to check landing page of your advertisers for malware and scams. They are already paying millions for this service.
Sometimes their system malfunctions (cough cough) and you get charged for those clicks but you fail to report them as fraudulent as you've no proof as the IPs and useragent all appear normal to ad agencies and advertisers.
So it's okay for organized crime to do it, but not normal plebs? The former is obviously a lot more harmful and in reality this is just a wakeup call to get their shit together.
People were dismissive of security far too long. I think it's actually good that AI instills some fear into people causing it to take a lot more seriously than they have previously done so.
Citation needed? I'm not a fan of residential proxies but these are pretty wild claims that aren't substantiated. Calling them a "national security threat" implies it is a threat to e.g., the continued existence of the United States. Is it really that? All of the activities described are already violations of the Computer Fraud and Abuse Act.
> At the very least, major American ISPs (Comcast, AT&T) should detect clearly suspicious activity coming from customer IPs and warn them to scan their computers, check their TV apps, and find whatever is turning their internet into a proxy.
What are we saying? Really, what are we saying? We should turn domestic ISPs into domestic surveillance apparatuses to "detect clearly suspicious activity"? What is "clearly suspicious activity"? Section 230 is still law.
The router could also be a good place to display live TCP usage and last 30 days usage and other network activity. There should also be a service that can identify whether specific IP addresses belong to botnets or other malicious infrastructure. If there are any abnormal activity mobile app can notify the user.
Governments could even purchase residential proxies themselves and check whether they're being used for malicious activity.
It isn't that difficult, but I guess governments just aren't that interested.
> The router could also be a good place to display live TCP usage and last 30 days usage and other network activity. There should also be a service that can identify whether specific IP addresses belong to botnets or other malicious infrastructure.
> It isn't that difficult, but I guess governments just aren't that interested.
I recently switched from my previous ISP because they randomly broke my ability to use my own router, and during the time when I had to default to using their own modem/router hardware before I could get the new ISP to come and set things up, I could barely even get a signal in my office upstairs (which had previously been connected via a mesh endpoint) because their router didn't expose any way for me to split 5 GHz and 2.4 GHz, and the router absolutely refused to let my devices connect via 2.4 GHz despite them having more than 90% packet loss due to the weak 5 GHz signal.
Regardless of how "easy" it is, I don't trust ISPs not to screw it up somehow and probably cause a lot more concrete damage (even if the individual issues they cause are smaller in magnitude) than the theoretical concerns of "national security" that, as far as I can tell from reading this thread, have caused a total of like a few hours of downtime one day in a couple decades.
If anonymous currencies can be considered a national security threat despite making up less than 0.001% of the world's total value, then why not anonymous packets?
Once the establishment has understood how to profit from anonymous currencies, then they'll be taken off the 'bad' list (no matter how many scams continue to be foisted upon the unwashed masses).
Disagree fairly strongly. IP addresses known to be malicious are unequal and should be treated as such.
Just because the idea of net neutrality exists doesn't mean it's true. I'm sure there's a specific context for it, and 'security' is not that context. It was about data/packet prioritisation wasn't it? Unrelated to security.
If so many sites and services didn't go out of there way to block or flag VPN users as suspicious then I would have no need to use residential / mobile proxies.
They're also a wholly necessary endeavor until the surveillance industry stops discriminating against IP ranges with endless captcha nagwalls. That, including its likely next development of remote attestation, is a much deeper problem for individual liberty. Individual liberty is itself more important than "national security" (which is more about protecting the government rather than the People) and thus needs to be addressed first.
Residential proxies are necessary to access services that try to ban VPNs without compromising anonymity. If they’re banned, it’s easier for websites to require information that can be used to track you (IP address), since most people don’t use VPNs they won’t care.
More importantly, a residential proxy can be mutual. If a group of people agree to forward traffic to each others’ IP, how is it not their business? It obfuscates their identities, but this is the Internet, not e.g. a government office or test center where you obviously can’t walk in with someone else’s ID.
Banning non-consensual proxies makes sense since those are effectively malware, even though they make anonymity harder, so does banning theft and here you’re stealing someone’s internet. I’m sure we can convince enough laymen to knowingly install proxies by paying them.
I think the fact that they're laymen means that there's still a pocket of non-consensuality. Taking advantage of someone's ignorance of potential consequences is as bad as malware in my opinion. If you're outlining all the potential bad outcomes before signing up a 'mark', then that's slightly more OK.
Myself being "not a layman" would definitely not allow anonymous usage of Internet that's tied to my home/residence/identity to be used by some internet rando just because they pay me money.
I would support, however, small family/friends groups who know and trust each other well enough to not do anything that'll cause a police raid on each others homes (having endured a police raid and the ensuing 8 months of not being told anything about the progress of whatever they're doing, it's not something I'd wish on a stranger, never mind a friend/family member).
No, your paranoia-outage is the true "national security threat" - a threat to freedom. Stop fanning the flames and calling for more government intervention.
They have their problems but how else am I supposed to scrape data from companies that want to hide it?
Apologies for the double post, but I've got an alternate perspective:
Do you allow the data you've scraped to be scraped? Do you share it as freely as you desire the 'companies that want to hide it' would?
I'm a bit on the fence about this, but leaning towards the "bad luck". I'm sure there's a large swathe of nuance that I'm missing, but my simplistic view is: If they don't want to be scraped then they don't get included in "the thing" which, at minimum, is a data point for consumers to make consumer decisions about.
It strongly depends on how scraped data is being used.
If your 'cat' hasn't been able to catch their 'mouse' then the cat needs to get smarter, or look for alternative sources of mice, or the cat should be considered 'unviable'.
Have you approached them to get access to their data? Have you explained to them how your service can benefit their business?
(I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect, but then I'm also an 'information wants to be free' kinda person, but the Internet is increasingly an untrustworthy place, so security is overruling narrative).
Take reddit for example. If you're not big tech they ignore you.
I'm kinda "fuck reddit".
That's cutting the tether from a _lot_ (there's no way to overstate this) of useful information, but that's potentially one of the great things about the open AI/LLM models, is that all that info is baked in there.
Disclaimer: As far as I understand it. Please educate me if I'm way off the mark.
Also, how much of the content of reddit is in Common Crawl? (same disclaimer applies to this comment)
I also understand 'the archival mindset', I hoard a bunch of data. But I also understand the logarithmic graph of futility.
I know there've been efforts to scrape information that MAGA is trying to purge from government websites.
Government website data should be openly available one way or another (at least within the country, and with reasonable security provision against obvious maliciousness).
If it has to be scraped then there may be other problems (which include lack of resources to make the data API accessible).
Exactly! I like the comment!
Anything connected to network can become a national security threat. IoT devices, internet connected smart appliances and all. Computers that run some OS is least of the worry because of the various layers of protection that can be added to them. But it is not the case with the IoT devices or SMART(!) devices where, if the firmware or any app ecosystem gets hacked / cracked / modified or sideloaded without any user intervention, could turn into a large zombie C2C botnet that can cause havoc.
Proxying is least of the worries!
When manufacturers bring out the smart features that require constant data collection, monitoring (obviously in the name of service quality, troubleshooting, firmware / app updates and subscription of services (!!)), it is oblivious to the fact that security and privacy at a personal / state / national level is never considered (for various reasons such as the design, profit interests and overheads / compliance perspectives and what not!) to be part of the ecosystem. Bad actors with sufficient knowledge and tools could exploit and profit from it.
It has become more of a planned obsolescence / profit / greed based industrial / corporate culture in rolling out unwanted stuff / monitoring and controlling as a feature that gets exploited to the core.
Honesntly, I don't have an answer for that, but have some thoughts that I wanted to share.
I do not want a cleaning robot or a water purifier or a printer or a TV or a refrigerator or an oven or a smart light bulb or a smart lock or even my car to have undisclosed, unwarranted connectivity to internet for whatever reasons.
I do not care if it is the manufacture or the service provider or whomsoever it may be!
I want all of the network connectivity options to be explained with all the controlling options and boundaries and data collection that happens on any of the connected medium to be as much transparent with the option of preventing or restricting what one does not want to go out of the device.
Will anyone do that?
It ill never happen! Sadly!
This problem will balloon further without adequate controls and killing the networking at a ground level would only be the only solution!
But, in the case of a connected device having an inbuilt connectivity option (like the m2m based options) in a smart vehicle, even that is not possible!
Ad networks like meta ads, google ads use residential IPs + small LLM to check landing page of your advertisers for malware and scams. They are already paying millions for this service.
Sometimes their system malfunctions (cough cough) and you get charged for those clicks but you fail to report them as fraudulent as you've no proof as the IPs and useragent all appear normal to ad agencies and advertisers.
It's an anti-consumer evil practice, but if it is a national threat that's on the nation's security infra.
somewhat related to @iammrpayments comment, the S in IoT stands for security.
The nation's security infra is a reflection of the nation's security legislation and regulation.
Aside: TheSInIoT is my DMZ's wifi password.
Does this take include, or exclude, DDOS?
First heard about this through this podcast. https://darknetdiaries.com/transcript/172/
That’s a scary read.
No they aren't.
> Actual data and identity loss of American citizens.
> Malware that can record video and audio from infected devices.
> Infrastructure for foreign covert influence campaigns.
> Botnets used in hacking and DoS attacks
These things have existed for 20+ years. Bad but not exactly a national security threat.
Mirai took out the internet in large parts of the US. The situation hasn't exactly improved since then.
Yea, but with AI we also have seen unprecedented amount of hackings etc.. using these residential proxies. Now even normal plebs can do so much harm.
So it's okay for organized crime to do it, but not normal plebs? The former is obviously a lot more harmful and in reality this is just a wakeup call to get their shit together.
People were dismissive of security far too long. I think it's actually good that AI instills some fear into people causing it to take a lot more seriously than they have previously done so.
Citation needed
Maybe IOT is the actual national security threat?
Citation needed? I'm not a fan of residential proxies but these are pretty wild claims that aren't substantiated. Calling them a "national security threat" implies it is a threat to e.g., the continued existence of the United States. Is it really that? All of the activities described are already violations of the Computer Fraud and Abuse Act.
> At the very least, major American ISPs (Comcast, AT&T) should detect clearly suspicious activity coming from customer IPs and warn them to scan their computers, check their TV apps, and find whatever is turning their internet into a proxy.
What are we saying? Really, what are we saying? We should turn domestic ISPs into domestic surveillance apparatuses to "detect clearly suspicious activity"? What is "clearly suspicious activity"? Section 230 is still law.
The router could also be a good place to display live TCP usage and last 30 days usage and other network activity. There should also be a service that can identify whether specific IP addresses belong to botnets or other malicious infrastructure. If there are any abnormal activity mobile app can notify the user.
Governments could even purchase residential proxies themselves and check whether they're being used for malicious activity.
It isn't that difficult, but I guess governments just aren't that interested.
> The router could also be a good place to display live TCP usage and last 30 days usage and other network activity. There should also be a service that can identify whether specific IP addresses belong to botnets or other malicious infrastructure.
> It isn't that difficult, but I guess governments just aren't that interested.
I recently switched from my previous ISP because they randomly broke my ability to use my own router, and during the time when I had to default to using their own modem/router hardware before I could get the new ISP to come and set things up, I could barely even get a signal in my office upstairs (which had previously been connected via a mesh endpoint) because their router didn't expose any way for me to split 5 GHz and 2.4 GHz, and the router absolutely refused to let my devices connect via 2.4 GHz despite them having more than 90% packet loss due to the weak 5 GHz signal.
Regardless of how "easy" it is, I don't trust ISPs not to screw it up somehow and probably cause a lot more concrete damage (even if the individual issues they cause are smaller in magnitude) than the theoretical concerns of "national security" that, as far as I can tell from reading this thread, have caused a total of like a few hours of downtime one day in a couple decades.
The optimum amount of residential proxies are non-zero.
Relevant Darknet Diaries episode: https://darknetdiaries.com/episode/172/
The interviewee seems pretty inept yet still managed to uncover something really interesting and nefarious.
Anonymous packets are a national security threat!
If anonymous currencies can be considered a national security threat despite making up less than 0.001% of the world's total value, then why not anonymous packets?
Once the establishment has understood how to profit from anonymous currencies, then they'll be taken off the 'bad' list (no matter how many scams continue to be foisted upon the unwashed masses).
Oh please. An IP address is an IP address. The whole idea that some IPs are more special than others flies in the face of net neutrality.
Disagree fairly strongly. IP addresses known to be malicious are unequal and should be treated as such.
Just because the idea of net neutrality exists doesn't mean it's true. I'm sure there's a specific context for it, and 'security' is not that context. It was about data/packet prioritisation wasn't it? Unrelated to security.
If so many sites and services didn't go out of there way to block or flag VPN users as suspicious then I would have no need to use residential / mobile proxies.
That’s a weird way to justify fucking over an uninvolved third party.
>fucking over an uninvolved third party
No one is being hurt by someone sharing their internet with me a few times a week.
They're also a wholly necessary endeavor until the surveillance industry stops discriminating against IP ranges with endless captcha nagwalls. That, including its likely next development of remote attestation, is a much deeper problem for individual liberty. Individual liberty is itself more important than "national security" (which is more about protecting the government rather than the People) and thus needs to be addressed first.
internet is a security threat, tell other...
Google propaganda. Botnet proxies, agreed. All residential proxies, dumb. Educate yourself: https://layer3intel.com/blog/measuring-the-netnut-takedown
Sounds like fear mongering to justify some kind of otherwise unpopular surveillance laws
Residential proxies are necessary to access services that try to ban VPNs without compromising anonymity. If they’re banned, it’s easier for websites to require information that can be used to track you (IP address), since most people don’t use VPNs they won’t care.
More importantly, a residential proxy can be mutual. If a group of people agree to forward traffic to each others’ IP, how is it not their business? It obfuscates their identities, but this is the Internet, not e.g. a government office or test center where you obviously can’t walk in with someone else’s ID.
Banning non-consensual proxies makes sense since those are effectively malware, even though they make anonymity harder, so does banning theft and here you’re stealing someone’s internet. I’m sure we can convince enough laymen to knowingly install proxies by paying them.
I think the fact that they're laymen means that there's still a pocket of non-consensuality. Taking advantage of someone's ignorance of potential consequences is as bad as malware in my opinion. If you're outlining all the potential bad outcomes before signing up a 'mark', then that's slightly more OK.
Myself being "not a layman" would definitely not allow anonymous usage of Internet that's tied to my home/residence/identity to be used by some internet rando just because they pay me money.
I would support, however, small family/friends groups who know and trust each other well enough to not do anything that'll cause a police raid on each others homes (having endured a police raid and the ensuing 8 months of not being told anything about the progress of whatever they're doing, it's not something I'd wish on a stranger, never mind a friend/family member).