Two obvious trends are cloud computing and mobile access. They are complementary. As the number of people and devices on the net increases, our ability to shape traffic on the demand side gets worse. Spikes in demand will happen faster and reach higher levels over time. Mobile devices exacerbate the demand side problems by greatly increasing both the number of people on the net and the fraction of their time they are able to access it.
Large traffic volumes both create and demand large data. Our tools for processing tera- and petabyte datasets will improve dramatically. Map/Reduce computing (a la Hadoop) has created attention and excitement in this space, but it is ultimately just one tool among many. We need better languages to help us think and express large data problems. In particular, we need a language that makes big data processing accessible to people with little background in statistics or algorithms.
Speaking of languages, many of the problems we face today cannot be solved inside a single language or application. The behavior of a web site today cannot be adequately explained or reasoned about just by examining the application code. Instead, a site picks up attributes of behavior from a multitude of sources: application code, web server configuration, edge caching servers, data grid servers, offline or asynchronous processing, machine learning elements, active network devices (such as application firewalls), and data stores. "Programming" as we would describe it today--coding application behavior in a request handler--defines a diminishing portion of the behavior. We lack tools or languages to express and reason about these distributed, extended, fragmented systems. Consequently, it is difficult to predict the functionality, performance, capacity, scalability, and availability of these systems.
Some of this will be mitigated naturally as application-specific functions disappear into tools and frameworks. Companies innovating at the leading edge of scalability today are doing things in application-specific behavior to compensate for deficiencies in tools and platforms. For example, caching servers could arguably disappear into storage engines and no-one would complain. In other words, don't count the database vendors out yet. You'll see key-value stores and in-memory data grid features popping up in relational databases any day now.
In general, it appears that Objects will diminish as a programming paradigm. Object-oriented programming will still exist... I'm not claiming "the death of objects" or something silly like that. However, OO will become just one more paradigm among several, rather than the dominant paradigm it has been for the last 15 years. "Object oriented" will no longer be synonymous with "good".
Some people have talked about "polyglot programming". I think this is a red herring. Polylgot is a reality, but it should not be a goal. That is, programmers should know many languages and paradigms, but deliberately mixing languages in a single application should be avoided. What I think we will find instead is mixing of paradigms, supported by a single primary language, with adjunct languages used only as needed for specialized functions. For example, an application written in Scala may mix OO, functional, and actor-based concepts, and it may have portions of behavior expressed in SQL and Javascript. Nevertheless, it will still primarily be a Scala application. The fact that Groovy, Scala, Clojure, and Java all run on Java Virtual Machine shouldn't mislead us into thinking that they are interchangeable... or even interoperable!
Regarding Java. I fear that Java will have to be abandoned to the "Enterprise Development" world. It will be relegated to the hands of cut-rate business coders bashing out their gray business applications for $30 / hour. We've passed the tipping point on this one. We used to joke that Java would be the next COBOL, but that doesn't seem as funny now that it's true. Java will continue to exist. Millions of lines of it will be written each year. It won't be the driver of innovation, though. As individual programmers, I'd recommend that you learn another language immediately and differentiate yourself from the hordes of low-skill, low-rent outsource coders that will service the mainstream Java consumer.
(...)To my great surprise, data storage has become a hotbed of innovation in the last few years. Some of this is driven by the high-scalability fetishists, which is probably the wrong reason for 98% of companies and teams. However, innovations around column stores, graph databases, and key-value stores offer developers new tools to reduce the impedance mismatch between their data storage and their programming language. We spent twenty years trying to squeeze objects into relational databases. Aside from the object databases, which were an early casualty of Oracle's ascension, we mostly focused on changing the application code through framework after framework and ORM after ORM. It's refreshing to see storage models that are easier to use and easier to modify.
(...)
Comentarios, discusiones, notas, sobre tendencias en el desarrollo de la tecnología informática, y la importancia de la calidad en la construcción de software.
Mostrando las entradas con la etiqueta ORM. Mostrar todas las entradas
Mostrando las entradas con la etiqueta ORM. Mostrar todas las entradas
domingo, octubre 31, 2010
Oyendo reflexiones sobre el software futuro
Mencionado por Jean Bezivin en Twitter: algunas reflexiones de Michael Nygard acerca del futuro del desarrollo de software, que suenan razonables, aunque quizá a Java le reste más (y mejor) recorrido que el que Michael estima. Lo fundamental, la computación móvil y su asociada, la "computación en la nube", aparecen forzando la marcha según sus requerimientos:
jueves, enero 11, 2007
Ralph Johnson sobre RDB y OODB
Johnson dedica una nota a una discusión iniciada en la lista Squeak-Dev sobre las ventajas y desventajas de las bases de datos relacionadas comparadas con las bases de datos orientadas a objeto, desde el punto de vista de Smalltalk. Toda la discusión es de interés, pero hay al menos dos participantes destacables, aparte de Ralph Johnson: Todd Blanchard, y Göran Krampe.
Qué dice Johnson acerca de cuándo una base de datos relacional es apropiada:
La discusión sigue...
Qué dice Johnson acerca de cuándo una base de datos relacional es apropiada:
Making a DBMS fast, reliable, etc. is hard. When you reuse a software component you are making your system depend on it and if your system is mission-critical then you do not want it to depend on software that is not reliable or is not well supported. One of the main advantages of RDBMSs is that they are old, boring technology that is produced by companies that have been selling them for a long time and that knows how to support them. And make money off them!Entre arquitectos basados en una concepción orientada a objetos (se trata de una lista relacionada con Smalltalk), el paradigma de las bases de datos relacionales no ajusta bien:
Goran pointed out that systems based on RDBMSs tend to grow unpredictably and become big balls of mud. But most systems tend to grow unpredictably and become big balls of mud! The usual argument for RDBMSs is that by making a good data model, you can have all sorts of applications reuse the data, and that the data will live much longer than the applications. As Goran pointed out, this would be nice if it happened, but in practice, the data is often highly coupled to the application.Con una posición pragmática, Johnson justifica al modelo relacional:
If you don't have good architects, big systems will end up as big balls of mud. A lot of companies live with it, but there are certainly big payoffs if you can avoid it. The main problem is that there aren't enough good architects to go around. One of the advantage of a RDBMS is that it is fairly easy to understand so below average programmers can still get systems running. Below average programmers will not build the best architectures, though. So, the fact that RDBMs lead to big balls of mud is actually a sign of an advantage. You can use a RDBMS when you have good architects and when you don't. You'll end up with completely different systems, but that is because of the quality of your architects, not because of the technology you use.Blanchard suma objeciones a las bases de datos orientadas a objeto:
(...) my favorite persistence pattern was the one of Prevayler, but that an experienced architect should know about XML, OODBMSs, and RDBMSs, because each had its place. In fact, there are at least two ways of using an RDBMS from an OO language. One is to use a object-relational mapping system, but another is to use SQL directly and to have the domain model in tables rather than the objects.
(...) One of the advantages of thinking of RDBMSs as a pattern is that it makes you stop asking "is it a good idea" and start asking "what are its advantages and disadvantages" and "when should we use it instead of its alternatives". For example, if speed is your number one criterion then you probably shouldn't use a RDBMS. That is why RDBMSs are big in enterprise applications and not in supercomputer applications. An enterprise app must be fast enough, but if you can process all your transactions every day and can give responses to people in a second or less, your app is probably fast enough. Enterprise developers often complain about performance, but compared to real-time developers or supercomputer programmers, performance is not very important to them, and the fact that RDBMSs are so popular among them is proof of that. On those occasions where performance IS important, get a new database.
One of the advantages of an RDBMS is that it tolerates bad data pretty well. It is usually not hard to look at your data and see what is wrong. XML is like this, too. OO databases tend not be like this, though Gemstone is an exception because you can use Smalltalk tools to inspect and modify the data.
OODBs are seductive. They are easy to get started with. For one thing, you don’t have to do a data model, just your object model. Your code is your model. You make objects, stick them in containers, and forget about them. Sounds great, right?...Y Göran Krampe recusa cada uno de estos argumentos, apoyándose en su experiencia con Gemstone.
But as anyone who has lived with an OODBMS for any period of time knows, Object databases are great, until they’re not, and then they truly suck. Here’s why:
1) Concurrency is very poor. (...)
2) Constant re-fetching of data every transaction makes keeping user interface elements up to date very expensive.(...)
3) Schema migration is hard, if not impossible. Your object defines your format.
4) Death by a trillion bug fixes.(...) [y el #6:6) Bugs are forever.]
5) OODBMS providers have limited resources and will only support versions up to one year old.(...)
7) No security.
8) Garbage Collection is not universally available. Orphaned junk is common.
9) No ad hoc query capability. You have to write a new program to view any data at all. You need to write programs to update reference data. You need a program to do anything at all with your data. No fixing problems with a quick line of SQL. Searching for unanticipated patterns is difficult. [Object Oriented Databases]
La discusión sigue...
jueves, enero 04, 2007
oodb en google groups
Un tema recurrente en el grupo comp.objects de Google Groups, es la discusión entre bases de datos relacionales y el diseño orientado a objetos. Aquí ya se han mencionado estas discusiones más de una vez. El centro de ésta discusión es el cuestionamiento de la capacidad del diseño orientado a objetos de resolver algunas áreas de problemas, donde las bases de datos relacionales tienen un papel principal. Existe en el grupo una corriente que, basada en la experiencia del desarrollo "procedural" y el uso de SQL, cuestiona la validez de OOD, generando debates de mucho interés, con participación de personas calificadas.
Algunos de ellos:
el último activo a hoy, Databases as Objects
Relational Databases & OO
Object Databases
Relational Databases Notation Question
...y otros.
Algunos de ellos:
el último activo a hoy, Databases as Objects
Relational Databases & OO
Object Databases
Relational Databases Notation Question
...y otros.
viernes, mayo 27, 2005
Abarca y Devora II - Descripción del invento
Data Structure Mappings: Un esquema Relacional a un esquema Orientado a Objetos; cómo es cubierto por la patente apuntada en la nota anterior.
Lo que sigue es la lista de Claims de la patente:
What is claimed is:
Lo que sigue es la lista de Claims de la patente:
What is claimed is:
1. A system that facilitates mapping arbitrary data models, comprising a mapping component that receives respective metadata from at least two arbitrary data models, and maps expressions between the data models.
2. The system of claim 1, the data models are query languages.
3. The system of claim 1, the data models are data access languages.
4. The system of claim 1, the data models are data manipulation languages.
5. The system of claim 1, the data models are data definition languages.
6. The system of claim 1, the data models include at least an object model and at least one relational model where the object model is mapped to at least one of the relational models.
7. The system of claim 1, the data models include object models where one of the object models is mapped to at least one of the other object models.
8. The system of claim 1, the data models include an XML model and at least one relational model where the XML model is mapped to at least one of the relational models.
9. The system of claim 1, the data models are XML models where one of the XML models is mapped to at least one of the other XML models.
10. The system of claim 1, the data models include an XML model and at least one object model where the XML model is mapped to at least one of the object models.
11. The system of claim 1, the data models are relational models where one relational model is mapped to at least one of the other relational models.
12. The system of claim 1, the data models include an XML model and an object model where the object model is mapped to at least one of the XML models.
13. The system of claim 1, the data models are of the same structure.
14. The system of claim 1, the mapping component receives topology data that is derived from the metadata.
15. The system of claim 1, the data models are read-only.
16. The system of claim 1, the data models are mapped without modifying the metadata or structure of the data models themselves.
17. The system of claim 1, the expressions comprise at least one of a structure, field, and relationship.
18. The system of claim 17, the expressions mapped between the data models are at least one of the same, different, and a combination of the same and different.
19. The system of claim 1, the mapping component creates structural transformations on the data of a data model by at least one of creating or collapsing hierarchies, moving attributes from one element to another, and introducing new relationships.
20. The system of claim 1, the mapping component relates and connects the same mappable concepts between the data models.
21. The system of claim 1, the mapping between the data models is directional.
22. The system of claim 1, the data models include a source domain and a target domain, such that a structure and field of the target domain can be mapped at least once.
23. The system of claim 1, the data models include a source domain and a target domain, such that a structure and field of the source domain can be mapped multiple times.
24. The system of claim 1, the data models include a source domain and a target domain such that the mapping component allows a user to operate on the data models through a query language of the target domain.
25. The system of claim 1, the data models include a source domain and a target domain such that the source domain is the persistent location of the data.
26. The system of claim 1, the data models include a source domain and a target domain such that mapping translates a query written in a query language of the target domain into a query language of the source domain.
27. The system of claim 1, the mapping component facilitates automatically synchronizing updates made in a target data model to a source data model.
28. The system of claim 1, the mapping component includes a mapping file that maps like concepts of the respective metadata.
29. A computer executing the system of claim 1.
30. A system that facilitates mapping arbitrary data models, comprising: source metadata that represents source concepts of a source data source; target metadata that represents target concepts of at least one target data source; and a mapping component that receives the source metadata and the target metadata and maps the concepts from the source data source to the target metadata associated with one or more of the target data sources.
31. The system of claim 30, the mapped concepts are at least one of the same, different, and a combination of the same and different.
32. The system of claim 30, the data sources are of the same structure.
33. The system of claim 30, the source concepts and the target concepts are the same in both of the data sources.
34. The system of claim 30, the mapping component relates and maps the same concepts between the data sources.
35. The system of claim 30, the source and target concepts include a relationship element that is a link and association between two structures in the same data source.
36. The system of claim 35, the relationship element defines how a first structure relates to a second structure in the same data source.
37. The system of claim 30, the source data source and the target data source are disposed on a network remote from the mapping component.
38. The system of claim 30, the mapping component is local to at least one of the source data source and the target data source.
39. A method of mapping data between data models, comprising: receiving respective metadata from at least two arbitrary data models; and mapping expressions between at least two of the data models based upon the metadata.
40. The method of claim 39, the expressions mapped between the two data models are the same expressions.
41. The method of claim 39, further comprising defining a source data schema and a target data schema and information missing in the schemas.
42. The method of claim 39, further comprising transforming data during a mapping of a source data model to a target data model using a function.
43. The method of claim 39, further comprising synchronizing changes made in a target data model with a source data model.
44. A method for mapping arbitrary data models, comprising: receiving source metadata that represents source concepts of a source data source and target metadata that represents target concepts of a target data source; and mapping the concepts between the source and target data sources based upon the source metadata and the target metadata.
45. The method of claim 44, the data sources are of the same structure.
46. The method of claim 44, further comprising, creating a variable in a source domain; restricting the variable with conditions; and mapping the variable to the target concept.
47. The method of claim 46, the variable is created at least one of implicitly and explicitly.
48. The method of claim 46, the variable represents an empty result set.
49. The method of claim 44, the mapping is stackable between the source and target data via one or more intermediate mapping stages.
50. The method of claim 44, further comprising selecting an optimal path for mapping between the source and the target.
51. The method of claim 50, the optimal path is selected with a central control entity based on at least one of available bandwidth and interruptions in the path.
52. The method of claim 44, further comprising accessing a mapping algorithm in response to selecting an optimal path between a plurality of the data sources and plurality of the data targets.
53. The method of claim 52, the mapping algorithm is associated with a structure of the source data and the target data.
54. A system that facilitates mapping data between arbitrary data models, comprising: means for receiving source metadata that represents source concepts of a source data source and target metadata that represents target concepts of a target data source; and means for mapping the concepts between the source and target data sources based upon the source metadata and the target metadata.
55. The system of claim 54, the means for mapping includes a mapping means that relates and maps the same concepts between the data sources.
jueves, mayo 26, 2005
Mel Brooks: "Abarca y Devora"
Una noticia de los últimos días: un grupo de desarrolladores patentó el mapeo entre estructuras de datos. Extractado por The Server Side:
Primer contacto con la noticia, una advertencia temprana de Jack Herrington en CGN-Talk, desde donde se puede leer el contenido completo de la patente aprobada, y los nombres de sus propietarios. Luego, la nota en The Server Side apuntada en el título de éste artículo, aribuyendo a Microsoft la acción, y, fundamentalmente, las consecuencias. Como en la discusión de The Server Side se dice, una patente no es definitiva, sino controversial; pero implica un avance definido hacia un objetivo restrictivo: poner una idea que en varios aspectos es de dominio público y académico, en manos de un grupo de beneficiarios de futuros reclamos de propiedad intelectual.
Aunque es temprano para atribuirle un padre, los "indicios vehementes" señalan uno, y uno acostumbrado a moverse con estos valores.
La siguiente es la introducción descriptiva del invento:
The patent is A data mapping architecture for mapping between two or more data sources without modifying the metadata or structure of the data sources themselves. Data mapping also supports updates. The architecture also supports at least the case where data sources that are being mapped, are given, their schemas predefined, and cannot be changed.En la lista de propietarios de la patente aparece un número importante de personas vinculadas a Microsoft. La voz común es considerar que es Microsoft mismo quien avanza sobre una idea que fue ampliamente aplicada a través de muchos años por muchas empresas, que puede incluír distintos perfiles de problemas, pero que es señalada especialmente como amenaza a los diseños existentes en OR mapping (mapeo de relacional a objeto y viceversa).
Primer contacto con la noticia, una advertencia temprana de Jack Herrington en CGN-Talk, desde donde se puede leer el contenido completo de la patente aprobada, y los nombres de sus propietarios. Luego, la nota en The Server Side apuntada en el título de éste artículo, aribuyendo a Microsoft la acción, y, fundamentalmente, las consecuencias. Como en la discusión de The Server Side se dice, una patente no es definitiva, sino controversial; pero implica un avance definido hacia un objetivo restrictivo: poner una idea que en varios aspectos es de dominio público y académico, en manos de un grupo de beneficiarios de futuros reclamos de propiedad intelectual.
Aunque es temprano para atribuirle un padre, los "indicios vehementes" señalan uno, y uno acostumbrado a moverse con estos valores.
La siguiente es la introducción descriptiva del invento:
[0004] The present invention disclosed and claimed herein, in one aspect thereof, comprises a mapping format designed to support a scenario where two (or more) data sources need to map to each other, without modifying the metadata or structure of the data sources themselves. Mapping is provided, for example, between an Object space and a relational database, Object space and XML data model, an XML data model and a Relational data model, or mapping could be provided between any other possible data model to XML, relational data model, or any other data model. The mapping format supports updates, and also supports the case where both data sources being mapped are given, their schemas are predefined, and cannot be changed (i.e., read-only). An approach that was previously used to map, for example, XML data to a relational database required making changes to the XML schema definition (XSD) file in order to add annotations. The mapping format of the present invention works as if the XSD file is owned by an external entity and cannot be changed. It also allows reuse of the same XSD file, without editing, for multiple mappings to different data sources (databases, etc.).Quisiera destacar el último párrafo de la introducción:
[0005] Each data model exposes at least one of three concepts (or expressions) to mapping: structure, field, and relationship. All of these concepts can be mapped between the data models. It is possible that one data model may have only one or two of the expressions to be mapped into another data model that has three expressions. The mapping structure is the base component of the mapping schema and serves as a container for related mapping fields. A field is a data model concept that holds typed data. Relationship is the link and association between two structures in the same data model, and describes how structures in the same domain relate to each other. The relationship is established through common fields of the two structures and/or a containment/reference where a structure contains another structure. These are just examples of relationships, since other relationships can be established (e.g., siblings, functions, . . . ). The present invention allows establishing arbitrary relationships. A member of a data model can be exposed as a different mapping concept depending on the mapping context.
[0006] Semantically, mapping is equivalent to a view (and a view is actually a query) with additional metadata, including reversibility hints and additional information about the two mapped domains. When one data source is mapped to another, what is really being requested is that it is desired that the Target schema is to be a view of the Source schema. Mapping is a view represented in one data domain on top of another data domain, and defines the view transformation itself. Mapping can create complex views with structural transformations on the data, which transformations create or collapse hierarchies, move attributes from one element to another, and introduce new relationships.
[0007] Mapping relates and connects the same mappable concepts between two or more mapped models. Mapping is also directional. Thus, one domain is classified as a Source and the other is classified as a Target. The directionality of mapping is important for mapping implementation and semantics, in that, a model that is mapped as a Source has different characteristics then a model that is mapped as a Target. The Target holds the view of the source model, where the mapping is materialized using the query language of the target domain. The Source is the persistent location of the data, and mapping translates the query written in the target domain query language to the source domain query language. One difference between Source and Target is that a structure or field from the Source or Target model has some restrictions regarding the number of mappings that can apply for structures and fields. In the target domain, a structure and a field can only be mapped once, whereas in the Source domain, a structure and a field can be mapped multiple times. For example, a Customers table can be mapped to a BuyingCustomer element and a ReferringCustomer element in the Target domain. However, a local element or local class can be mapped only once. Another difference that stems for the directional attribute of mapping is that mapping allows users to operate on the mapped models through the query language of the target domain (e.g., using XQuery for mapping a Relational model to an XML model, and OPath, for mapping a Relational model to an Object model).
[0008] Another important attribute of the mapping architecture is that of being updateable. In the past, developers had to write code to propagate and synchronize changes between the domains. However, in accordance with the present invention, the mapping engine now performs these tasks. That is, when the user creates, deletes, or modifies a structure in the target domain, these changes are automatically synchronized (persisted) to the source domain by the target API and mapping engine.
[0009] In another aspect thereof, the mapping architecture is stackable where multiple stages of mappings may occur from a source to a target.
[0010] To the accomplishment of the foregoing and related ends, certain illustrative aspects of the invention are described herein in connection with the following description and the annexed drawings. These aspects are indicative, however, of but a few of the various ways in which the principles of the invention may be employed and the present invention is intended to include all such aspects and their equivalents. Other advantages and novel features of the invention may become apparent from the following detailed description of the invention when considered in conjunction with the drawings....
Suscribirse a:
Entradas (Atom)