HN user

mtrunkat

21 karma
Posts2
Comments23
View on HN

In our case (Apify.com) there was a complete outage of SQS (15mins+), most likely DNS problems + EC2 instances got restarted probably as a result of an SQS outage.

EDIT: Also AWS Lambda seems to be down and AWS EC2 APIs having a very high error rate and machines slow startup times.

Apify | Full-time | On-site or Remote | Prague, Czech republic

Join us on our journey in building a web scraping and automation platform to help both the world’s biggest companies and upcoming startups leverage the web's full potential. Each month, we save people tons of hours of manual work by turning more than a billion web pages into data and automating over 15 million tasks on the web.

-----------

Our tech stack:

* Frontend: React.js, styled-components, Storybook, Cypress

* Backend: TypeScript/Node.js, Next.js, Express.js, Nest.js, Jest

* Infra: AWS, Kubernetes, Helm, MongoDB, Redis, DynamoDB, S3, ...

* Monitoring: New Relic, LogDNA, Sentry, PagerDuty

* Tools: GitHub, ZenHub, Notion, GSuite

* Process: two-week sprints, code reviews, tests, automating whatever we can, deploying multiple times per day

-----------

Backend engineer (NodeJS) https://apify.com/jobs#backend-engineer-(node.js)

Frontend engineer (React) https://apify.com/jobs#frontend-engineer-(react)

Web automation developer (NodeJS) https://apify.com/jobs#web-automation-developer-(node.js)

-----------

More positions here: https://apify.com/jobs

Apify | Full-time | On-site or Remote | Prague, Czech republic

Apify runs on a highly scalable infrastructure that runs millions of web automation and web scraping jobs weekly to process way over a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 to process our workloads, multiple Kubernetes clusters to host our applications and services, and store data in MongoDB, DynamoDB, S3, Redis, and SQS.

Our system is built with Node.js and many other technologies. We use Kubernetes, Helm, Github Actions, and CloudFormation to make our deployments smooth, and ship to production multiple times per day. We monitor using New Relic One, Sentry, LogDNA, and other services.

We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day and we are proud of. We love open-source and contribute to it–check out our Github https://github.com/apify. Help us make the Web more open and programmable!

DevOps/SRE https://apify.com/jobs#devops/sre-(site-reliability-engineer...

NodeJS Backend engineer https://apify.com/jobs#backend-engineer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site | Prague, Czech republic

Apify runs on a highly scalable infrastructure that runs millions of web automation and web scraping jobs weekly to process way over a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 to process our workloads, multiple Kubernetes clusters to host our applications and services, and store data in MongoDB, DynamoDB, S3, Redis, and SQS.

Our system is built with Node.js and many other technologies. We use Kubernetes, Helm, Github Actions, and CloudFormation to make our deployments smooth, and ship to production multiple times per day. We monitor using New Relic One, Sentry, LogDNA, and other services.

We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day and we are proud of. We love open-source and contribute to it–check out our Github https://github.com/apify. Help us make the Web more open and programmable!

DevOps/SRE https://apify.com/jobs#devops/sre-(site-reliability-engineer...

NodeJS Backend engineer https://apify.com/jobs#backend-engineer

More positions here: https://apify.com/jobs

This article describes some of the techniques and MongoDB Cloud features we used to debug performance issues and expose sub-optimal queries.

Apify | Full-time | On-site | Prague, Czech republic

Apify runs on a highly scalable infrastructure that runs millions of web automation and web scraping jobs weekly to process way over a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 to process our workloads, multiple Kubernetes clusters to host our applications and services, and store data in MongoDB, DynamoDB, S3, Redis, and SQS.

Our system is built with Node.js and many other technologies. We use Kubernetes, Helm, Github Actions, and CloudFormation to make our deployments smooth, and ship to production multiple times per day. We monitor using New Relic One, Sentry, LogDNA, and other services.

We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day and we are proud of. We love open-source and contribute to it–check out our Github https://github.com/apify. Help us make the Web more open and programmable!

DevOps/SRE https://apify.com/jobs#devops/sre-(site-reliability-engineer...

NodeJS Backend engineer https://apify.com/jobs#backend-engineer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site | Prague, Czech republic Apify runs on a highly scalable infrastructure that runs millions of jobs weekly to process way over a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 to process our workloads, multiple Kubernetes clusters to host our applications and services, and store data in MongoDB, DynamoDB, S3, Redis, and SQS.

Our system is built with Node.js and many other technologies. We use Kubernetes, Helm, Github Actions, and CloudFormation to make our deployments smooth, and ship to production multiple times per day. We monitor using New Relic One, Sentry, LogDNA, and other services.

We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day and we are proud of. We love open-source and contribute to it–check out our Github https://github.com/apify. Help us make the Web more open and programmable!

DevOps/SRE (Site reliability engineer) https://apify.com/jobs#devops/sre-(site-reliability-engineer...

NodeJS Backend engineer https://apify.com/jobs#backend-engineer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site | Prague, Czech republic

Apify runs on a highly scalable infrastructure that runs millions of jobs weekly to process way over a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 to process our workloads, multiple Kubernetes clusters to host our applications and services, and store data in MongoDB, DynamoDB, S3, Redis, and SQS.

Our system is built with Node.js and many other technologies. We use Kubernetes, Helm, Github Actions, and CloudFormation to make our deployments smooth, and ship to production multiple times per day. We monitor using New Relic One, Sentry, LogDNA, and other services.

We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day and we are proud of. We love open-source and contribute to it–check out our Github https://github.com/apify. Help us make the Web more open and programmable!

DevOps/SRE (Site reliability engineer) https://apify.com/jobs#devops/sre-(site-reliability-engineer...

NodeJS Backend engineer https://apify.com/jobs#backend-engineer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site | Prague, Czech republic

Apify runs on a highly scalable infrastructure that runs millions of jobs weekly to process over a billion web pages every month . We run on a cluster of Linux servers on Amazon EC2 to process our workloads, multiple Kubernetes clusters to host our applications and services, and store data in MongoDB, DynamoDB, S3, Redis, and SQS.

Our system is built with Node.js and many other technologies. We use Kubernetes, Helm, Github Actions, and CloudFormation to make our deployments smooth, and ship to production multiple times per day. We monitor using New Relic One, Sentry, LogDNA, and other services.

We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day and we are proud of. We love open-source and contribute to it–check out our Github https://github.com/apify. Help us make the Web more open and programmable!

DevOps/SRE (Site reliability engineer) https://apify.com/jobs#devops/sre-(site-reliability-engineer...

NodeJS Backend engineer https://apify.com/jobs#backend-engineer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site | Prague, Czech republic

Apify builds software technology and infrastructure that helps small startups and the world’s biggest companies leverage the full potential of the Web—the largest source of information ever created in the history of humankind. To help achieve this, we've developed a serverless cloud platform and tools focused on web scraping and automation.

Apify runs an infrastructure that processes over a billion web pages every month. We use: Node.js, MongoDB, AWS (EC2, S3, DynamoDB, SQS, ECS, Lambda, ECR, EKS ...), Linux, Kubernetes, ReactJS, Next.js, headless Chrome / Puppeteer, ...

Backend Engineer https://apify.com/jobs#backend-engineer

Fullstack Engineer https://apify.com/jobs#fullstack-engineer

Frontend Engineer https://apify.com/jobs#frontend-developer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site/Remote | Prague | Czech republic

Apify builds software technology and infrastructure that helps small startups and the world’s biggest companies leverage the full potential of the Web—the largest source of information ever created in the history of humankind. To help achieve this, we've developed a serverless cloud platform and tools focused on web scraping and automation.

Apify runs an infrastructure that processes over a billion web pages every month. We use: Node.js, MongoDB, AWS (EC2, S3, DynamoDB, SQS, ECS, Lambda, ECR, EKS ...), Linux, Kubernetes, ReactJS, Next.js, headless Chrome / Puppeteer, ...

We're always looking for talented people to join the family, regardless of open positions. If you like what we do and are up for the challenge, get in touch at jobs@apify.com

Remote Open-source Node.js Engineer https://apify.com/jobs#open-source-engineer-node-js

On-site Backend or Full-stack Engineer https://apify.com/jobs#fullstack-engineer

More positions here: https://apify.com/jobs

Apify | Full-time | On-site/Remote | Prague | Czech republic

Apify builds software technology and infrastructure that helps small startups and the world’s biggest companies leverage the full potential of the Web—the largest source of information ever created in the history of humankind. To help achieve this, we've developed a serverless cloud platform and tools focused on web scraping and automation.

Apify runs an infrastructure that processes over a billion web pages every month. We use: Node.js, MongoDB, AWS (EC2, S3, DynamoDB, SQS, ECS, Lambda, ECR, EKS ...), Linux, Kubernetes, ReactJS, Next.js, headless Chrome / Puppeteer, ...

Remote Open-source Node.js Engineer https://apify.com/jobs#open-source-engineer-node-js

On-site Full-stack Engineer https://apify.com/jobs#fullstack-engineer

On-site Front-end Engineer https://apify.com/jobs#frontend-developer

More positions here: https://apify.com/jobs

Apify.com | Product Manager | Full-time | On-site | Prague, Czech republic

Apify builds software technology and infrastructure that helps small startups and the world’s biggest companies leverage the full potential of the web—the largest source of information ever created in the history of humankind. To help achieve this, we've developed a serverless cloud platform and tools focused on web scraping and automation. We are looking for a product manager who will work with a customer success team, uncover customer needs and then work with the development team on the delivery of a truly excellent solution.

Who are we looking for?

- You know how to research user requirements, analyze them, prioritize and formulate a development road map

- You are able to take over the prioritization of our product development and move it to the next level

- Solving people's problems (both your users and your teams) drives you

- You are able to work closely with a variety of people and lead a team (through influence, not authority)

- You have an ability to thrive in a fast-paced, collaborative, agile environment

- Technical background is a big plus

We offer

- Full-time job in Prague, Czech Republic (we have office in Lucerna Palace)

- Friendly, inspiring and no-bullshit work environment

- You'll work with some of the most talented and experienced developers in Prague

- Flexible working hours, possibility to work remotely and nobody counts holidays, as long as the work gets done

- Stock options, free lunches, unlimited supply of coffee and beer

https://apify.com/jobs

Apify.com | Full-stack engineer | Full-time | On-site | Prague, Czech republic

Apify builds software technology and infrastructure that helps small startups and the world’s biggest companies leverage the full potential of the web—the largest source of information ever created in the history of humankind.

Apify runs an infrastructure that processes almost a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 and store data in MongoDB, DynamoDB, S3, Redis and SQS. The system is built with Node.js and React. Apify actors run in Docker, and inside them runs Apify SDK, headless Chrome with Puppeteer, PhantomJS, or pretty much anything. We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use every day.

Who are we looking for?

- You have experience building backend and frontend systems

- You are highly skilled at developing and debugging in JavaScript/Node.js, or have this skill in some other programming language and are able to learn JavaScript quickly

- You are familiar with Linux

- You are able to speak and write in English

- Your knowledge of any technologies mentioned above is a plus

We offer

- Full-time job in Prague, Czech Republic (we have office in Lucerna Palace)

- Friendly, inspiring and no-bullshit work environment

- You'll work with some of the most talented and experienced developers in Prague

- Flexible working hours, possibility to work remotely and nobody counts holidays, as long as the work gets done

- Stock options, free lunches, unlimited supply of coffee and beer

https://apify.com/jobs

Apify.com | Full-stack engineer | Full-time | On-site | Prague, Czech republic

Apify runs on a highly-scalable infrastructure that processes almost a billion web pages every month. We run on a cluster of Linux servers on Amazon EC2 and store data in MongoDB, DynamoDB, S3, Redis and SQS. The system is built with Node.js, Meteor.js and React. Apify actors run in Docker, and inside them runs Apify SDK, headless Chrome with Puppeteer, PhantomJS, or pretty much anything. We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day.

Who are we looking for?

- You have experience building backend and frontend systems

- You are highly skilled at developing and debugging in JavaScript/Node.js, or have this skill in some other programming language and are able to learn JavaScript quickly

- You are familiar with Linux

- You are able to speak and write in English

- Your knowledge of any technologies mentioned above is a plus

- A university degree in software engineering or computer science is a big plus

We offer

- Full-time job in Prague, Czech Republic (we have office in Lucerna Palace)

- Friendly, inspiring and no-bullshit work environment

- You'll work with some of the most talented and experienced developers in Prague

- Flexible working hours, possibility to work remotely and nobody counts holidays, as long as the work gets done

- Stock options, free lunches, unlimited supply of coffee and beer

https://apify.com/jobs

Apify | Infrastructure engineer | Prague, Czechia | ONSITE

Apify runs on a highly-scalable infrastructure that processes almost billion web pages every month. We run on a cluster of Linux servers on Amazon EC2, store data in MongoDB, DynamoDB, S3, Redis and SQS, and use LogDNA, Newrelic and CloudWatch for monitoring. The core system is built with Node.js and Apify actors run in Docker. We're passionate about delivering the best service to our customers using the best technology possible. Apify is made by developers for developers. We're building a product that we use ourselves every day.

We're looking for experienced engineers who know how to design and setup scalable distributed computing systems and who are able to learn quickly and work independently. You will be helping us improve all parts of the Apify platform and building our current and future products. Join our team and help us make the web more programmable!

-----

Who are we looking for?

- You have experience with AWS, GCP or some other public cloud

- You have experience building backend infrastructure and know some of the technologies mentioned above

- You know Linux inside out

- You are skilled at developing and debugging in Node.js, or have this skill in some other programming language and are willing to learn Node.js

- Experience with Docker, Kubernetes or other container technology is a plus

- You are able to speak and write in English

-----

https://apify.com/jobs

Apify | Infrastructure engineer | Prague, Czechia | ONSITE

Apify runs on a highly-scalable infrastructure that enables it to load and analyze millions of web pages every day. We employ a cluster of Linux servers running on Amazon EC2 and store data in MongoDB, DynamoDB, S3, Redis and SQS. The whole software stack is based on JavaScript, we're using Node.js for backend services along with Meteor and React for the frontend. Actors are running inside our custom Docker container orchestrator and web scraping tasks are performed using headless Chrome, Puppeteer, PhantomJS, Selenium or any other suitable tool. We're passionate about delivering the best service to our customers using the best technology possible, constantly improving all parts of our system. Apify is made by developers for developers. We're building a product that we ourselves use every day.

We're looking for experienced engineers who know how to design and build scalable systems and who are able to learn quickly and work independently. You will be helping us improve all parts of the Apify platform and building our current and future products. Join our team and help us make the web more programmable!

-----

Who are we looking for?

- You have experience with AWS, GCP or some other public cloud

- You have experience building backend infrastructure and know some of the technologies mentioned above

- You know the Linux ecosystem inside and out

- You are skilled at developing and debugging in JavaScript/Node.js, or have this skill in some other programming language and are willing to learn JavaScript

- Experience with Docker or other container technology is a plus

- Experience with Kubernetes is a plus

- You are able to speak and write in English

-----

https://apify.com/jobs

We have found 2 related issues:

- Sometimes Promise returned by page.close() never resolves so it's good to call Promise.race() on that together with a Promise that resolves after some timeout period (30s?)

- Sometimes Chrome process doesn't get killed so we are also manually killing remaining Chrome process after browser.close()

Problem with fixed concurrency is that required memory per Chrome process varies a lot.

We (https://www.apify.com) are solving this by autoscaling number of parallel Puppeteer instances based on memory and CPU. Our open source SDK (http://github.com/apifytech/apify-js) implements this using class PuppeteerCrawler (https://www.apify.com/docs/sdk/apify-runtime-js/latest#Puppe...) which internally uses AutoscaledPool that provides autoscaling:

- https://github.com/apifytech/apify-js/blob/master/src/autosc...

- https://www.apify.com/docs/sdk/apify-runtime-js/latest#Autos...

Sadly this feature is currently limited to our platform because it's mainly build for running in Docker containers. And as Docker container don't know about it's CPU consumption vs limits it requires a notifications about reaching its CPU limit from underlying platform. We have solved this using Websocket events. But we are currently working on extending this to work anywhere/locally.