Two obvious trends are cloud computing and mobile access. They are complementary. As the number of people and devices on the net increases, our ability to shape traffic on the demand side gets worse. Spikes in demand will happen faster and reach higher levels over time. Mobile devices exacerbate the demand side problems by greatly increasing both the number of people on the net and the fraction of their time they are able to access it.
Large traffic volumes both create and demand large data. Our tools for processing tera- and petabyte datasets will improve dramatically. Map/Reduce computing (a la Hadoop) has created attention and excitement in this space, but it is ultimately just one tool among many. We need better languages to help us think and express large data problems. In particular, we need a language that makes big data processing accessible to people with little background in statistics or algorithms.
Speaking of languages, many of the problems we face today cannot be solved inside a single language or application. The behavior of a web site today cannot be adequately explained or reasoned about just by examining the application code. Instead, a site picks up attributes of behavior from a multitude of sources: application code, web server configuration, edge caching servers, data grid servers, offline or asynchronous processing, machine learning elements, active network devices (such as application firewalls), and data stores. "Programming" as we would describe it today--coding application behavior in a request handler--defines a diminishing portion of the behavior. We lack tools or languages to express and reason about these distributed, extended, fragmented systems. Consequently, it is difficult to predict the functionality, performance, capacity, scalability, and availability of these systems.
Some of this will be mitigated naturally as application-specific functions disappear into tools and frameworks. Companies innovating at the leading edge of scalability today are doing things in application-specific behavior to compensate for deficiencies in tools and platforms. For example, caching servers could arguably disappear into storage engines and no-one would complain. In other words, don't count the database vendors out yet. You'll see key-value stores and in-memory data grid features popping up in relational databases any day now.
In general, it appears that Objects will diminish as a programming paradigm. Object-oriented programming will still exist... I'm not claiming "the death of objects" or something silly like that. However, OO will become just one more paradigm among several, rather than the dominant paradigm it has been for the last 15 years. "Object oriented" will no longer be synonymous with "good".
Some people have talked about "polyglot programming". I think this is a red herring. Polylgot is a reality, but it should not be a goal. That is, programmers should know many languages and paradigms, but deliberately mixing languages in a single application should be avoided. What I think we will find instead is mixing of paradigms, supported by a single primary language, with adjunct languages used only as needed for specialized functions. For example, an application written in Scala may mix OO, functional, and actor-based concepts, and it may have portions of behavior expressed in SQL and Javascript. Nevertheless, it will still primarily be a Scala application. The fact that Groovy, Scala, Clojure, and Java all run on Java Virtual Machine shouldn't mislead us into thinking that they are interchangeable... or even interoperable!
Regarding Java. I fear that Java will have to be abandoned to the "Enterprise Development" world. It will be relegated to the hands of cut-rate business coders bashing out their gray business applications for $30 / hour. We've passed the tipping point on this one. We used to joke that Java would be the next COBOL, but that doesn't seem as funny now that it's true. Java will continue to exist. Millions of lines of it will be written each year. It won't be the driver of innovation, though. As individual programmers, I'd recommend that you learn another language immediately and differentiate yourself from the hordes of low-skill, low-rent outsource coders that will service the mainstream Java consumer.
(...)To my great surprise, data storage has become a hotbed of innovation in the last few years. Some of this is driven by the high-scalability fetishists, which is probably the wrong reason for 98% of companies and teams. However, innovations around column stores, graph databases, and key-value stores offer developers new tools to reduce the impedance mismatch between their data storage and their programming language. We spent twenty years trying to squeeze objects into relational databases. Aside from the object databases, which were an early casualty of Oracle's ascension, we mostly focused on changing the application code through framework after framework and ORM after ORM. It's refreshing to see storage models that are easier to use and easier to modify.
(...)
Comentarios, discusiones, notas, sobre tendencias en el desarrollo de la tecnología informática, y la importancia de la calidad en la construcción de software.
domingo, octubre 31, 2010
Oyendo reflexiones sobre el software futuro
martes, mayo 08, 2007
La normalización relacional y OOD
Responding to V4vijayakumar...I really hope this was not a homework question... B-)
> How object-oriented design can be normalized? Normalization, anywayThe Class Model in UML is underlain by the same relational data model
> related to OOD?
branch of set theory that underlies DBMSes. As a result the Class Model
needs to normalized to Third Normal Form just like an RDB schema.
[Normal Forms above third are rarely relevant because they mostly deal
with identifier conventions and we rarely use explicit identifiers for
objects.]Most OOA/D authors don't talk explicitly about normalization. However, every OOA/D author will provide a suite of rules for constructing Class Models that essentially ensure Third Normal Form under the guise of things like one-fact-one-place. For example, the most comprehensive book available on Class Modeling is Leon Starr's "Executable UML: How to Build Class Models". Leon doesn't even mention Normal Form as far as I recall, yet he provides the most comprehensive set for guidelines for normalization that I have seen.
In a nutshell we have:
1NF: all responsibilities must be a simple domain. For knowledge attributes this means that the attribute must be described in terms of an abstract data type (ADT) that can be manipulated as if it were a scalar. For attributes that can be expressed in terms of fundamental values, the domain of data values must have a single semantics. So a domain of {UNSPECIFIED, 5, 6, 7} is invalid because it captures two separate semantics: valid data values of 5, 6, 7 and whether or not the data is specified at all.
For behaviors this means that the behavior responsibility must be cohesive and self-contained. Self-contained means it can depend on knowledge attributes but it can't depend upon other object's behavior responsibilities. (Note that this comes for free if one follows the methodology's dictums about encapsulation.)
2NF; all responsibilities are fully dependent on the object identity. Typically objects do not have explicit identity attributes but they do have an unambiguous mapping to some some uniquely identifiable problem space entity. This means the "value" of the property depends solely on what problem space entity is abstracted in the object.
As a practical matter 2NF is not very relevant to OO development because it is really about compound identifiers (i.e., multiple attributes combine to define the object identity). What 2NF is saying is that if there are multiple explicit identity attributes, then the "value" of a non-identity attribute must be dependent on /all/ of the identity attributes, not just some of them. A classic example of this is:
[Housing Development]
+ developmentID // identifier
...[Subdivision]
+ developmentID // identifier
+ subdivisionID // identifier
...[House]
+ developmentID // identifier
+ subdivisionID // identifier
+ houseID // identifier
+ style
+ builderThe style attribute is clearly dependent on the particular House identity, which must be fully specified. The same thing seems true for the 'builder' attribute since each House is built by one builder. But suppose construction policy is that a builder builds all the houses in a particular subdivision. Now the 'builder' value is fully specified if one only knows {developmentID, subdivisionID}. So the 'builder' attribute really belongs in the [Subdivision] class. [Note that if the development id seriously homogenized, all Houses in the same subdivision might have the same style. In that case, style also belongs in
[Subdivision].]
3NF: all responsibilities depend upon nothing but the object identity. Essentially this means that the "value" of a responsibility cannot depend upon knowledge attributes that are not explicit identity attributes. A classic example of this problem is:
[House]
+ address // identifier
+ builder
+ style
+ cost
...The problem here is that it is highly unlikely that 'cost' is only dependent on the House identity. In fact, it is probably dependent on the style or on the combination of {builder, style}. IOW, only the /combination/ of {builder, style, and cost} is dependent solely on the identity of House, not the individual values. So, assuming cost is solely dependent on style, we need:
[House]
+ address
+ builder
+ style // referential attribute
...[Style]
+ style
+ costwhere the unique combination of values is captured indirectly through
the relationship to [Style].---
One must be careful not to confuse coincidental values or data domains with dependency. Consider Washing Machine and Refrigerator objects that both have a 'color' attribute. If the colors are designed to be color coordinated from the same manufacturer, they will have identical data domains for 'color'. It is quite possible that an object from both sets may be colored chartreuse. Nonetheless they are quite different things. How is that the 'color' attribute doesn't violate 1NF (same data domains semantics) and 3NF (both have the same color)?
The trick is to think for such generic qualities in terms of 'color of'. IOW, the color of a Washing Machine is chartreuse and the color of a Refrigerator is chartreuse. Thus the color of a Washing Machine is not semantically the same as the color of a Refrigerator even though the value is the same. Similarly, we have:
[Appliance]
+ color
A
|
+--------+-------------+
| |
[Refrigerator] [Washing Machine]The notion of the color of an Appliance is something shared; it raises the level of abstraction of color of a Refrigerator and color of a Washing Machine to a common ground for both. Since a Refrigerator is an Appliance, it has the color attribute.
This example, though, underscores an important difference between
normalization applied to OO Class Models and the Data Models used for RDB Schemas. In the OO case Refrigerator can implement a different data domain than a Washing Machine for the 'color' attribute (i.e., they aren't required to be color coordinated). That is not true in Data Models where there will be exactly one data domain for 'color'.The reason is that in Data Modeling the [Appliance] table is instantiated separately from the subclass tables and the subclass tables do not have a 'color' attribute. So there is only one attribute in one table with one data domain.
In contrast, in the OO Class Model the superclasses cannot be instantiated separately so an object resolves the entire tree. Thus the object identity /includes/ the [Appliance] properties. In addition, the Class Model only identifies the responsibility (i.e., What it is); it does not define its implementation (i.e., How it does it).
So Refrigerator and Washing Machine can provide different implementations of the 'color' attribute, such as different data domains. IOW, we resolve object identity at the leaf level of the tree through inheritance. Thus unique data domains can be associated with subclasses even though the responsibility is identified in a superclass.