Web scraping services for public data at scale
Prices, listings, directories, search results, company data. I build scrapers and data pipelines that collect public data reliably, clean it, and deliver it where you need it. I work with companies that have no developer of their own, and with teams that need a scraper that keeps running.
I'm all in on AI, so the work goes fast. I stay careful with data quality and with the rules of the sites involved.
I usually reply within one working day.
What I build
From a single scraper to a pipeline that runs on its own.
Scrapers for one source
Prices, listings, product data or directory entries from a site you need to watch. The data arrives structured, as CSV, JSON or straight in your database.
Large-scale collection
Many sources, millions of pages, queues and several servers. My own platforms collect data at this size.
Headless browsers
Sites that load their content with JavaScript need a full browser. I use headless browsers where simple requests are not enough.
Proxies and reliability
Retries, proxy rotation, rate limits, and alerts when a site changes its layout.
Cleaning and matching
I remove duplicates, clean up fields and match records to your own data. AI helps with messy text, like pulling prices or sizes out of descriptions.
Pipelines and delivery
Scheduled runs, change detection, and delivery to your database, a spreadsheet, an API or your app's admin area.
How I work
You don't need to have everything figured out before we talk. An idea, access to the repo, a screen recording or a list of problems is enough. I look at what is there, tell you what I would do and what it costs, and then I build it.
I use Claude Code, Codex and my own agent workflows for reading code, planning changes, writing tests, refactoring and debugging. It is not vibe coding. It is senior engineering with a much faster loop.
AI-first, on public data
Claude Code and Codex write parsers fast and fix them when a site changes. I decide what gets collected and check the data. I stick to public data, the sites' terms and the law.
I stay in control
AI writes a lot of my code. I still make every technical decision, read every change before it ships, and test it. Fast, but meticulous: the speed is only worth something if what goes live is right.
The way you want to work
Some clients want to look at every change before it goes live, others want me to deploy. Some want a weekly call, most prefer written updates. I work well asynchronously, and I still ask the important questions when they come up.
I think like an owner
I still run my own online products. So when I build or fix yours, I think about traffic, conversion, revenue and support as much as about the code, and I tell you when something is not worth building.
Ways to work with me
Fixed price, a flat monthly rate, or hourly. It's always me doing the work, not an agency. If you're not sure, describe what you need and I'll suggest one.
Examples of my work
Scraping and data pipelines in my own platforms and client products.
Millions of opportunities a month
A Laravel platform on more than 30 servers that finds websites, discovers contact details and runs outreach on its own. Its scraping engine uses headless browsers and proxy rotation.
Search results for AI agents
Scraping of search engines and other web sources as input for AI agent workflows, running on queues with Redis and Horizon.
Data pipelines from TMDb
Pipelines that pull movie and TV data from TMDb, watch for new releases and changes, and feed an AI content engine.
What clients say
"What stood out about working with Vincent was how methodical and efficient he was."
"You have a chat, you express what you need, and he comes back with what you asked for."
"Vincent is one of those rare developers who can jump into a complex codebase, become useful immediately, and solve real problems without hand-holding."
Questions
Is web scraping legal?
Collecting public data is allowed in many cases, but it depends on the data, the site's terms and the country. Personal data needs extra care under GDPR. I tell you where I see a problem. For a legal answer you need a lawyer.
What happens when the site changes?
Scrapers break when a layout changes. I add checks that notice missing data, and then I fix the scraper.
How do we get the data?
As CSV or JSON files, in your database, through an API, or in an admin area where your team can search it.
Can you scrape sites that need a login?
Only with an account you are allowed to use for this, and only when the site's terms permit it.
Which language do you write scrapers in?
Mostly Laravel and PHP. I use Python or Node.js when a library there fits better.
How do you charge?
Fixed price for a job with a clear outcome, a flat monthly rate for ongoing work, or hourly. Describe what you need and I'll suggest one.
Get a quote
Which sites, which data, how often, and where it should end up. A few sentences are enough. I'll reply with what I would do and what it would cost.
If you would rather talk first, book a 30-minute call or email contact@vincentschmalbach.com.