let's imagine we want to build a url shortening service something like tinyurl or bitly that allows us to take a longer url and create a shortcut that we can more easily share with other people online all right so before we jump into designing the api or the architecture for this application let's think about some of the core features that we're going to want to support in our system so we know that a high level that we want to support things like creating a shortcut for a url and we know that we want users or visitors
who visit that shortcut to be redirected to the original url so i'm going to say redirect to original url we want these shortcuts to be unique for users when they create them so we're not sort of reusing or overwriting old shortcuts and we want to support a large number of shortcuts so there's sort of some technical goals here as well around having a high volume of requests and we want our system to be performant and reliable so those are some properties that will sort of inform how we design our api as well as how we
design the architecture of our application here so with that in mind let's think about some of the api endpoints that we're likely going to need to support these features so let's jump over here and let's write out some api endpoints so i know at a high level we're we're just going to need a create shortcut url endpoint here that's going to take the original url and it's going to then return a shortcut url which will include sort of a random shortcut id which we'll talk a little bit about implementing in a little while we're also
going to want to have a way to visit this url and then retrieve the url that we're going to be redirected to and so i'm just going to call this um you know get original url and this will take in the shortcut url okay now that we've laid out some of the api endpoints and features that we want to support let's talk about the overall architecture of our application and the components that we're going to need to be able to build so let's talk about both of these use cases where we have a user who's
creating a shortcut and a user who needs to visit a shortcut and be redirected we can have both of these users contact our api server on the back end and this will help either create a url or redirect the user so let's talk about this user who's creating a shortcut first they're going to call our shortcut url endpoint and then our api server is going to insert a new entry into our database to store this url shortcut like so and in the case where a user is visiting a shortcut we're going to do the same
thing except we're just going to look up the shortcut from our database and then return a 302 redirect back to the user so let's imagine that they're visiting short.com xyz after our server looks up that shortcode if it exists it will return a 302 to the destination url now a 302 is a standard http status code that indicates that a user should be redirected so the browser would know how to interpret this and would forward the user straight on to their destination all right so now let's talk a little bit about how we're going to
store these shortcuts in our database and how we're going to generate these shortcut urls in the first place so with that in mind let's let's talk about the database table that we're going to use here now i think this is a pretty simple sort of key value store so we could pick any sort of database that would work for this it could be a nosql database something like dynamodb or it could even be a relational database as long as we can scale it to support a large number of users either by sharding the database using
the shortcut url as a partition key or building a or just using a platform that automatically gives us this sort of scalability and and partitioning out of the box so we're basically going to have a shortcuts table and it's going to have a few properties that are important to us one is this shortcut id which is going to be our our generated string we're going to have our url destination which is also going to be a string and then we might want to store something like the user id who created the shortcut in the first
place so that people can see and update their url shortcuts now when it comes to generating these shortcut ids we need to think about a few properties of our application and our goals here so we know that it's going to depend on two things one is going to be the number of possible urls we want to support that'll help determine the length of the string and number of combinations we need to support and then we also need to think about things like do we want our url to be short and easy to remember or do
we want it to be you know random and cryptographically secure so that it's hard to guess now i think in this case we're going to want to support a relatively short and easy to remember code and we might want to support something like up to a billion possible urls so with that in mind we can design a url generation scheme that will match these requirements so i think in order to support a billion url combinations we're going to have to do a little bit of math here but all we need to do is sort of
think about the possible number of choices we have for each character let's stick to a standard alphabet with 26 characters and then we'll raise it to some power that is equal to the length of our string in this case let's pick like six or seven because that'll give us billions of possible combinations here so that's all we have all we need to do and then in order to actually pick this url shortcut we just need to sort of use an out of the box random generation function or we could use something more complex like taking
a unique hash of the url and the current timestamp for example but that'll give us what we need and then we can just store this in our database and retrieve it later now let's talk a little bit about how we can make our system even more performant and scalable so one thing to consider is the different types of request patterns we're going to be receiving from these different types of users so for example we might expect the system to be relatively read heavy because there's probably going to be many more users visiting shortcut urls than
there are users creating shortcut urls so as a result we're going to be doing a lot of reads from our database and not as many writes so with that in mind it might be worth thinking about implementing some sort of caching policy here and adding a cache to our system so and the reason for this is we want our reads to be really performant and fast especially for url shortcuts that are highly trafficked so we might expect spikes and traffic visiting a particular url and we want those to be really fast now so i'm going
to go ahead and pencil in a cache into our system diagram here and the idea is that when we visit a shortcut url we'll go ahead and put it into our cache so that subsequent reads can uh just access the cache first and so this would sort of reduce our latency here from you know anywhere from 100 milliseconds to perhaps sub-millisecond for an in-memory cache and then we'll use something like an lru cache policy so that we're only storing the url shortcuts that are being frequently or recently visited and that will make our system much
faster now that basically covers the simple implementation of a url shortening service that let's make things a little bit more complex and say that we want to add some analytics information so that users know how often people are visiting their url shortcuts and can even see usage patterns over time now in order to do that it's worth considering how our system might change because we're talking about going from a very read-heavy system to a very right-heavy system so now every time someone visits a url we also need to write some information back to our database
to store this usage information so one way we could do this and sort of reduce the bottleneck that might be caused by this level of rights going back to our database would be to actually every time someone visits a url we increment account also in our caching service and then with some level of granularity we can sort of flesh flush out the counts from our cache and then store them in our database in sort of a batch operation so that would be one way to do it but either way we need to store some of
this information in our database or in some sort of separate analytics database and it's worth thinking about how we would structure this table to sort of capture this information in a in a time series way so that we can calculate these reports so i'm going to call this a new table called visits and it's going to have the same shortcut url which is a string and then we're also going to have a count which would be an integer and this would represent the number of visits we got in a particular time series so we're not
going to start every single visit here we're going to store sort of a bucket of counts in a particular time range depending on the granularity of reporting that we want and then we're also going to store sort of the date or time at which this was stored which would be a time stamp so this might be a particular minute or a particular day depending on what kind of reporting we wanted and that would basically be it and then in order to generate reports what we would do is we could visualize the number of visits to
this url over time we could also generate summary statistics like the number of visits per month by summing over this table or creating a roll-up table that we also keep updated over time and with that that basically covers our system so now we've talked about how to build a basic url shortener we've talked about how to build in some time series analytics and we've talked about strategies like caching that improve the performance of our system