Scalability is the ability to adjust resources to meet demand. If more people use your application, a scalable system can add capacity so it stays responsive; if fewer people use it, you can remove capacity so you stop paying for it. On-premises, scaling means buying and installing hardware, which takes weeks. In the cloud, scaling is a configuration change, a command or a rule, and it takes minutes. Scalability matters in both directions: adding capacity protects performance, and removing it protects your budget, because in a consumption-based model idle capacity is still billed. There are two directions of scaling. Vertical scaling, also called scaling up, gives an existing resource more power, such as moving a virtual machine to a size with more CPU cores or memory, or moving a database to a higher performance tier. Scaling down is the reverse. Vertical scaling is simple, because the application still runs on one machine and needs no redesign, but there is a ceiling on how large a single machine can be, and resizing a VM usually requires a restart, which means a short interruption.
Keep reading for free
Create a free StudyToCert account to read the rest of this lesson: 7 more sections, 6 key terms, a real-world example, an exam tip and self-check questions. Every lesson, lab and practice test is free with an account.